OpenAI releases 722 AI-written math papers on open problems, but not the prompts mathematicians asked for

White OpenAI logo over a dark background, for a story on OpenAI releasing 722 AI-written math papers

Key points

  • OpenAI posted a large batch of AI-written math papers
  • An advisory group asked labs for their prompts a week earlier
  • Some proofs haven't been checked by computer yet

OpenAI published 722 mathematical manuscripts on October 6, grouped into 372 families of related results, which it says were produced by an unreleased AI model.

The release comes a week after the Advisory Group on Mathematics and Artificial Intelligence, a panel of mathematicians hosted at the Institute for Advanced Study, asked AI labs to publish the prompts behind results like these and to stop testing hard problems on private models.

Why mathematicians are asking for more transparency

On September 21, OpenAI said a model it began training on August 28 had resolved the Navier-Stokes problem, one of the seven $1 million Millennium Prize problems, and "more than 100 long-standing open problems across most areas of mathematics." The same post said "the pace of its progress in mathematics has surprised the mathematicians within OpenAI."

That post cited an open letter from mathematicians, "A Severe Misalignment of AI in Mathematics." It also named nine initial members of the advisory group, including Fields Medal winners Timothy Gowers, Martin Hairer, and Edward Witten. OpenAI wrote that the group "will not be responsible for advising us on how to pace our internal progress on mathematics."

The group published its recommendations on September 29, after collecting more than 600 replies from mathematicians. For results that the people who prompted the AI don't yet understand, it asked labs to name the model and to publish the prompts, a summary of the reasoning, and the computing cost.

"We strongly recommend that AI labs refrain from treating the release of mathematical results as marketing vehicles to promote their models," the group wrote. It also said results should go into repositories that aren't "controlled by any AI lab."

OpenAI's October 6 post said it had consulted the group and drawn on its advice. It published the results on its own GitHub account and said it is "continuing to explore other community-hosted alternatives." It released 10 reasoning summaries, compute estimates, and attempt statistics, and said it will fund workshops and conferences on the results.

An OpenAI spokesperson told Scientific American that the team is taking the group's guidelines seriously and doing its best to comply, but that the company isn't bound by them. The magazine reported that OpenAI shared the average compute time and some statistics, with no prompts.

How the results were produced and checked

The model was given about 4,000 problems over the course of the evaluation, according to the repository's notes. OpenAI estimates that each result used computing resources equivalent to roughly three hours of ChatGPT Pro thinking, on average.

OpenAI said the catalog came from grouping that output and "requiring an appropriate level of significance." Exceptions to the standard process include a zero-free region for the Riemann zeta function and a proof of the Hodge conjecture for a class of geometric objects called CM abelian varieties. The zeta function write-up was also "human edited for readability."

The release includes summaries of the model's reasoning for 10 results. One covers the irrationality exponent of pi, a measure of how closely fractions can approximate the number.

Many of the proofs have versions written in Lean, a programming language that lets a computer check a proof's logic, but not all of them do yet. "Some of the unformalized results could have issues," the repository says. "We will endeavor to fix any such issues quickly."

An earlier batch of 10 results that OpenAI announced on August 1 has already drawn outside review. An audit posted to arXiv by Mikołaj Sienicki and Krzysztof Sienicki found "no confirmed substantive mathematical error in a principal result" in the reviews it examined, "although review depth varies." That audit covered the August results, not the manuscripts released in October.

How mathematicians reacted

The spokesperson told the magazine that almost every result came from a single prompt given to a single AI agent, though some may have taken multiple attempts.

"Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified," Andrew Sutherland, a mathematician at the Massachusetts Institute of Technology, told Scientific American.

Daniel Litt, a mathematician at the University of Toronto, said that if mathematicians want the answers, he sees no reason to ask the company "to keep them secret from us." "To me, it's going to be a good thing for mathematics," Litt said.

"We should ask for receipts," Sutherland said.

Frequently asked questions

What did OpenAI release on October 6, 2026?

OpenAI published 722 mathematical manuscripts, grouped into 372 families of related results, in a public GitHub repository. It says the results were produced by an internal AI model that hasn't been released. Many of the proofs come with versions written in Lean, a language that lets a computer check a proof, but not all of them do yet.

How much computing did each result use?

OpenAI estimates that each result used computing resources equivalent to roughly three hours of ChatGPT Pro thinking, on average. The model was given about 4,000 problems over the course of the evaluation.

What did the mathematics advisory group ask AI labs to do?

In recommendations published on September 29, 2026, the Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study, asked AI labs to stop testing advanced math problems on proprietary models. For results that the people who prompted the AI don't yet understand, it asked labs to name the model and publish the prompts, a summary of the reasoning, and the computing cost. It also asked labs not to treat math results as marketing for their models.

Did OpenAI release the prompts behind the results?

Scientific American reported that OpenAI shared the average compute time and some statistics, with no prompts. OpenAI did release summaries of the model's reasoning for 10 results. An OpenAI spokesperson told the magazine the company is taking the group's guidelines seriously but isn't bound by them.

Has anyone checked OpenAI's earlier math results?

An audit posted to arXiv reviewed 10 results OpenAI announced on August 1, 2026. It found no confirmed substantive mathematical error in a principal result in the reviews it examined, although review depth varied. That audit didn't cover the October manuscripts.

More coverage

Dennis Singleton
Dennis Singleton

Dennis Singleton was born in Australia and later moved to the United States. He has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.