
OpenAI dumps 372 AI-generated math proofs on GitHub, telling the academic world to keep up
Quick Answer
OpenAI has released 372 AI-generated mathematical proofs on GitHub, challenging traditional academic publishing.
Quick Take
The proofs, derived from a single AI model prompt, aim to address open problems, including advancements related to the Riemann hypothesis, while formal verification in Lean seeks to ease review bottlenecks.
Key Points
- The proofs include significant advancements in algorithms and the Riemann hypothesis.
- OpenAI's model produced results with an average compute cost of three hours per proof.
- Formalizations in Lean aim to verify logical correctness of the AI-generated proofs.
- The math community is divided on the implications of mass-produced proofs.
- OpenAI plans to fund workshops to enhance understanding of AI-generated results.
📖 Reader Mode
~3 min readOpenAI has released a large collection of mathematical results produced by an internal AI model. The proofs are published on GitHub rather than in academic journals, with formal verification intended to make review more practical.
OpenAI has published 372 new mathematical results generated by an internal frontier model. Each result is supposed to solve an open problem or make substantial progress toward one. The collection includes improvements to major computer algorithms and advances related to the Riemann hypothesis.
The company is hosting the results in a GitHub repository, complete with revision logs and citations. According to OpenAI, the same model already produced a solution to a Navier-Stokes problem that has been under formal review for weeks.
Nearly every result came from a single prompt to a single AI agent, OpenAI says, though some took multiple attempts. That's a sharp contrast to the Navier-Stokes solution, which required a swarm of 10,000 agents and millions of dollars in compute. On average, each result consumed roughly three hours' worth of ChatGPT Pro Thinking compute.
Formal verification could ease the review bottleneck
Many of the proofs come with formalizations in Lean, a programming language built for machine-checkable mathematical proofs. More formalizations are planned. The reason is practical: the sheer volume of AI-generated results could easily overwhelm the math community's capacity for manual review.
OpenAI also published details on its methodology, including summaries of the reasoning process, statistics on how many problems the model attempted, and estimates of compute costs.
Traditional journals aren't built for this pace
OpenAI put its results on GitHub instead of peer-reviewed journals. It's a statement move, since it implies that the traditional process of doing science is too slow for this volume of potentially new knowledge.
OpenAI consulted with the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and loosely followed their public recommendations, since the company didn't publish any prompts and only shared average compute costs rather than per-problem figures. The advisory group includes Fields Medal winner Timothy Gowers among other renowned mathematicians.
The company also announced plans to fund workshops and conferences focused on understanding AI-produced results. OpenAI acknowledged it wants to improve the quality of its citations and presentation, and says it's working on a responsible release of the model to "directly empower scientists with state-of-the-art capabilities."
OpenAI set one significant boundary for the advisory group beforehand, though: the mathematicians can advise on how results get communicated, but not on whether or how fast they're produced.
The math community is split on what mass-produced proofs actually mean
Lean formalizations can verify logical correctness, but they can't judge whether a result is mathematically relevant or original. OpenAI is betting that its results push the boundary of human knowledge. Whether the mathematical community agrees remains an open question. So far, reactions range from excitement to frustration.
In a recent open letter titled "A Severe Misalignment of AI in Mathematics," 25 Fields Medal winners warned of a deep disconnect between the AI industry's goals and those of mathematics. Problem-solving, they wrote, is merely a tool and proxy for the real goal of conceptual understanding and insight. Mass-producing true statements could destroy fertile ground rather than bring new ideas to life, they argued. The effect would spill over into other fields.
Gowers has warned that within one to two decades, mathematical literature could grow enormously while no human community remains that truly understands it. Fields Medal winner Terence Tao has added that training young mathematicians needs to emphasize the human side and tightly limit AI tool use so that genuine learning and understanding survive.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

