The 10 results span an unusually broad set of disciplines: high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics .
OpenAI did not release the Astra model itself. Instead, it released Lean 4 proof certificates on GitHub under an Apache 2.0 license, with a reported zero "sorry" count — meaning every step across all ten formalized proofs is fully verified . This approach sidesteps the usual debate over whether AI-generated reasoning is correct: the proof assistant enforces correctness mechanically, regardless of how the argument was discovered.
Even as researchers like Noam Brown called the announcement a major step for scientific reasoning , skeptics raised significant questions.
Not the big ones. Several observers noted that none of the 10 problems are Clay Millennium Prize problems (e.g., P vs. NP, Riemann Hypothesis) . The problems are serious open questions — the non-sofic group construction, for instance, has been open since Mikhail Gromov introduced soficity in 1999
— but they do not carry the same headline weight.
No peer review. The results were announced directly by OpenAI without prior publication in a peer-reviewed mathematics journal or conference proceedings . Critics argued that the line between AI-generated discovery and human-assisted framing remains blurry, and OpenAI itself acknowledged this tension by referencing the Leiden Declaration on AI and Mathematics
.
Generalization skepticism. Cognitive scientist Gary Marcus argued that success on formal math is a "special case" that may not generalize to broader scientific reasoning, and that expertise in one domain does not guarantee expertise across domains . Math lends itself to cheap, verifiable synthetic data, which is not true of most scientific fields.
Reproducibility concerns. Since the model is unreleased and the specific prompting methodology was not fully disclosed, independent verification of the generation process is impossible . Some critics focused on the difficulty of replicating the results without access to Astra itself.
Commercial timing. The announcement came as OpenAI is positioning its next major model family, raising questions about whether the research was timed for marketing impact rather than scientific rigor .
If the proofs survive expert scrutiny, several results would close questions open for decades. The headline achievement — the first explicit construction of a non-sofic group — resolves a question that has stood for more than 25 years . Even critics acknowledge that OpenAI supplied far more inspectable material than a typical corporate announcement
.
But the broader question is not whether Astra can do math. It is whether solving formal, verifiable problems in a narrow domain translates into general-purpose scientific reasoning. The answer, for now, is: not yet proven.