Sharing AI progress in mathematics
Quick Answer
OpenAI has released new mathematical results from its internal frontier model in a GitHub repository, including formal proofs in Lean and detailed transparency about the model's reasoning and compute usage.
Quick Take
The initiative aims to enhance collaboration with the math community and support further advancements in mathematics through workshops and conferences.
Key Points
- Results published in a GitHub repository with protocols for revisions and citations.
- Formal proofs shared in Lean, a programming language for checking mathematical proofs.
- Average compute usage for results equated to three hours of ChatGPT Pro thinking.
- OpenAI plans to fund workshops and conferences to promote understanding of AI-generated results.
- Commitment to improve future releases based on community feedback and best practices.
📖 Reader Mode
~2 min readWe’re releasing a broad range of new mathematical results produced by an internal frontier model.
As we look to improve how we share results with the math community, we’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study(opens in a new window) to develop best practices, and we have drawn on their advice and public recommendations(opens in a new window) to inform how we release these results.
For this release, we’re publishing the results in a GitHub repository, with protocols for paper revisions and citations. We’re continuing to explore other community-hosted alternatives for this release which meet the committee’s guidelines. For future releases, we are committed to further improving the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding.
As part of our GitHub repository, we are sharing formalizations of many of the proofs in Lean, a programming language that allows mathematical proofs to be checked by a computer. We will update the repository with more formalizations as we obtain them.
To promote scientific transparency and openness, we are also publishing additional details about how we obtained the results in the repository. These include 10 summaries of the model’s reasoning, estimations of compute spent in terms of Pro usage on ChatGPT, and statistics about the number of attempted problems. The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.
We want this progress to push the frontier of human knowledge and enable further progress in mathematics. We will be funding a series of workshops, conferences, and special programs around the understanding of major results produced by AI—we will share more on this in the near future.
We want to directly empower scientists with state-of-the-art capabilities and are working to responsibly release the model that produced these results. This is why it is important to continue to evaluate our internal frontier models on mathematics and other sciences, so we can accelerate developing the tools to advance those fields. We will continue to act on feedback from the community and update our standards for future disclosures of major scientific advancements.
— Originally published at openai.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from OpenAI Blog
See more →Scientific computing in the age of agentic AI
AI agents are transforming scientific computing by streamlining software development, enabling researchers to focus on discovery. Projects using Codex and Claude Code report accelerated development and improved maintenance, though challenges in validating AI outputs remain. Long-term stewardship of research software is crucial to ensure reliability and reproducibility.