Key points
- OpenAI posted a large batch of AI-written math papers
- An advisory group asked labs for their prompts a week earlier
- Some proofs haven't been checked by computer yet
OpenAI published 722 mathematical manuscripts on October 6, grouped into 372 families of related results, which it says were produced by an unreleased AI model.
The release comes a week after the Advisory Group on Mathematics and Artificial Intelligence, a panel of mathematicians hosted at the Institute for Advanced Study, asked AI labs to publish the prompts behind results like these and to stop testing hard problems on private models.
Why mathematicians are asking for more transparency
On September 21, OpenAI said a model it began training on August 28 had resolved the Navier-Stokes problem, one of the seven $1 million Millennium Prize problems, and "more than 100 long-standing open problems across most areas of mathematics." The same post said "the pace of its progress in mathematics has surprised the mathematicians within OpenAI."
That post cited an open letter from mathematicians, "A Severe Misalignment of AI in Mathematics." It also named nine initial members of the advisory group, including Fields Medal winners Timothy Gowers, Martin Hairer, and Edward Witten. OpenAI wrote that the group "will not be responsible for advising us on how to pace our internal progress on mathematics."
The group published its recommendations on September 29, after collecting more than 600 replies from mathematicians. For results that the people who prompted the AI don't yet understand, it asked labs to name the model and to publish the prompts, a summary of the reasoning, and the computing cost.
"We strongly recommend that AI labs refrain from treating the release of mathematical results as marketing vehicles to promote their models," the group wrote. It also said results should go into repositories that aren't "controlled by any AI lab."
OpenAI's October 6 post said it had consulted the group and drawn on its advice. It published the results on its own GitHub account and said it is "continuing to explore other community-hosted alternatives." It released 10 reasoning summaries, compute estimates, and attempt statistics, and said it will fund workshops and conferences on the results.
An OpenAI spokesperson told Scientific American that the team is taking the group's guidelines seriously and doing its best to comply, but that the company isn't bound by them. The magazine reported that OpenAI shared the average compute time and some statistics, with no prompts.
How the results were produced and checked
The model was given about 4,000 problems over the course of the evaluation, according to the repository's notes. OpenAI estimates that each result used computing resources equivalent to roughly three hours of ChatGPT Pro thinking, on average.
OpenAI said the catalog came from grouping that output and "requiring an appropriate level of significance." Exceptions to the standard process include a zero-free region for the Riemann zeta function and a proof of the Hodge conjecture for a class of geometric objects called CM abelian varieties. The zeta function write-up was also "human edited for readability."
The release includes summaries of the model's reasoning for 10 results. One covers the irrationality exponent of pi, a measure of how closely fractions can approximate the number.
Many of the proofs have versions written in Lean, a programming language that lets a computer check a proof's logic, but not all of them do yet. "Some of the unformalized results could have issues," the repository says. "We will endeavor to fix any such issues quickly."
An earlier batch of 10 results that OpenAI announced on August 1 has already drawn outside review. An audit posted to arXiv by Mikołaj Sienicki and Krzysztof Sienicki found "no confirmed substantive mathematical error in a principal result" in the reviews it examined, "although review depth varies." That audit covered the August results, not the manuscripts released in October.
How mathematicians reacted
The spokesperson told the magazine that almost every result came from a single prompt given to a single AI agent, though some may have taken multiple attempts.
"Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified," Andrew Sutherland, a mathematician at the Massachusetts Institute of Technology, told Scientific American.
Daniel Litt, a mathematician at the University of Toronto, said that if mathematicians want the answers, he sees no reason to ask the company "to keep them secret from us." "To me, it's going to be a good thing for mathematics," Litt said.
"We should ask for receipts," Sutherland said.



