Research#Model training#Content quality
OpenAI publishes 722 AI math papers
The openai/math GitHub release carries Apache-2.0 manuscripts with Lean proofs, over IAS advisory-group objections.

OpenAI published 722 mathematics manuscripts in the openai/math GitHub repository (Apache-2.0) on October 6, organized into 372 problem families. The author is an unnamed internal frontier model — trained from August 28 and described as significantly more capable than GPT-6 Astra — that worked through roughly 4,000 problems at an average cost of about three ChatGPT-Pro-equivalent hours of compute per result. Most manuscripts carry Lean formalizations; two results took different paths: a zero-free region for the Riemann zeta function at Re(s) > 11/12 (human-edited writeup) and a proof of the Hodge Conjecture for CM abelian varieties. The Verge and the New York Times covered the release the same day. GitHub counts differ by outlet — 722 manuscripts is the repository figure; some coverage cites 372 “families” as the headline number — because the organization maps many manuscripts onto single problems.
Key points
- Scale: 722 manuscripts / 372 problem families, Apache-2.0
- Model: unnamed internal frontier model (trained from Aug 28; stronger than GPT-6 Astra; first result Sept 5)
- Compute: ~3 ChatGPT-Pro-equivalent hours per result on average
- Verification: Lean formalizations for most; the README concedes unformalized results “could have issues,” promising versioned corrections
- Highlights: zeta zero-free region at Re(s) > 11/12; Hodge for CM abelian varieties; sample problems include the irrationality exponent of π, the Mahler conjectures, free group factor isomorphism
- Opposition: the IAS Advisory Group on Mathematics and AI (Sept 29) stated it does “not endorse this practice” and asked labs to stop testing advanced mathematics on proprietary models
The objections and the gap
The IAS group’s recommendations — issued after 600+ community replies — were specific: deposit results in repositories labs do not control, disclose the model name, prompts, summarized chain of thought, compute cost, and failed attempts. Against that checklist, the release scores partially: the repo is GitHub rather than an OpenAI property, and Lean plus ten abridged reasoning summaries provide real verifiability — but the model is unnamed, failures undisclosed, and the compute figure is an average. Half-executed is the honest description of where this governance tug-of-war stands.
How the math community digests it
Between the advisory group’s “stop” and OpenAI’s 722 papers sits the attention economy of working mathematicians: as The Verge and Unite.AI both note, the share of genuinely new theorems in the batch is undisclosed — much of it may be new proofs or strengthenings of known results. For professionals, panning 722 manuscripts is expensive; but the Lean-formalized subset can be triaged programmatically, and mathematicians on X are already organizing “Lean-first” sorting. The camps are polarized between the September endorsers and the advisory group’s caution, and no community consensus has formed.
Formalization is the moat
The line between “AI-written mathematical text” and “AI-proved mathematics” is Lean: machine-checkable proofs do not depend on a reviewer’s goodwill. The manuscripts that pass Lean are hard results; those that do not — which the README itself flags — re-enter the world of traditional review. That is also why OpenAI can release 722 at once: with formalization as the backstop, errors become versioned corrections, and review cost migrates from human reading to machine checking.
The collision with arXiv’s rate limit
arXiv’s new two-papers-per-month cap was built for human flooding; 722 manuscripts of AI capacity drive straight into the capacity ceiling of human review. The advisory group’s “stop” carries no enforcement power, but it frames the question the publication system must now answer: when proofs can be machine-checked, what is left for the journal to add? For researchers, the release is a free problem bank and comparison set; for publishers, it is the second time this month that AI productivity has backed the system into a corner — first the scale of 722 itself, then the increasingly hard-to-refute claim that machine-checkable proofs need no journal. The compromise taking shape: AI-generated results go to formal-verification repositories, human narratives to traditional publishing, and the two systems issue DOIs side by side.