OpenAI withdrew three mathematics manuscripts on Oct. 7, less than a day after publishing 722 AI-generated manuscripts on GitHub. A sign error in one paper invalidated its central argument and undermined two papers that depended on it. The episode is a reminder that making a mathematical manuscript public is not the same as independently verifying its claims.
24
35
What OpenAI released
On Oct. 6, OpenAI published manuscripts produced by an internal model that it had not released publicly. The repository grouped the work into 372 families: sets of related papers that can include a principal result, companion arguments, consequences or alternative proofs. The collection spans areas including algebra, geometry and computer science.
1
35
38
The family structure matters: 722 manuscripts did not mean 722 independent results. Some papers relied on arguments or constructions developed in other papers in the same family. That connection became central to the withdrawals.
1
24
The sign error that broke three papers
The problem appeared in Algebraicity of Weil classes on split abelian eightfolds. In a key argument, the paper assigned a geometric operation the sign +1 where, under the paper’s conventions, it should have been −1. That mistake broke a cancellation argument used to support the paper’s construction.
21
24
27
Two other manuscripts relied on that work: Algebraicity of Kuga–Satake Correspondences for K3 Surfaces and The rational Hodge conjecture for products of K3 surfaces. Once the underlying argument failed, their dependent claims could no longer be supported by it, so OpenAI withdrew all three. This was a failure of the papers’ arguments—not a disproof of the conjectures they addressed.
21
24
35
What changed in the repository
The Oct. 7 update reported the three withdrawals, revisions to 14 other manuscripts and reference updates in 13 companion papers. The catalog moved from 722 manuscripts to 719 while remaining organized into 372 families.
1
24
28
The repository and OpenAI’s release materials describe results at different stages of verification and warn that some may contain errors. OpenAI said its approach to sharing the work drew on consultation with the Advisory Group on Mathematics and Artificial Intelligence. That context supports treating the manuscripts as material for scrutiny, rather than as a set of settled discoveries.
1
33
39
What Lean formalization can—and can’t—tell readers
The collection includes Lean formalizations for some results. Lean is a proof assistant that lets computers check formalized mathematical proofs, but the available reports do not use one consistent measure of coverage: one summary counts about 300 of 719 top-line results as formalized, while another counts 162 of 722 manuscripts with a formalization of the main result. Those figures use different units, so neither should be read as meaning that every claim in a corresponding paper has been independently validated.
25
33
36
OpenAI’s consultation with the advisory group and the presence of formalized proofs provide context for evaluating the release; neither removes the need to check the written arguments and the scope of each claim. The withdrawals show why review of linked papers matters: an error in one shared argument can affect more than one manuscript.
24
33
39
The takeaway: treat the remaining claims as research to verify
The withdrawals establish that a specific argument failed and that two dependent papers could not stand on that argument. They do not establish that the other manuscripts are wrong—or that their claims are correct. Each result still needs assessment on its own merits, with attention to the proof, its assumptions and any dependencies on related work.
1
21
35