
OpenAI’s math solutions aren’t meeting the field’s standards yet
Quick Answer
OpenAI's recent math solutions, while released with input from elite mathematicians, still fail to meet community standards, particularly in ensuring human understanding.
Quick Take
A new paper highlights discrepancies between natural language proofs and formal expressions, raising concerns about the reliability of AI-generated solutions without human oversight.
Key Points
- Only 10 out of 719 proofs included the model's reasoning process.
- 42% of OpenAI's proofs lacked formalization as recommended by the advisory group.
- Discrepancies were found between natural language proofs and Lean code in OpenAI's solutions.
- Advisory group urges OpenAI to fund human mathematicians for meaningful validation.
- Mathematicians stress the importance of human engagement in understanding new results.
DeepSignal Analysis
What happened
OpenAI's recent release of solutions to complex math problems has not met the standards set by an advisory group of elite mathematicians. Despite some adherence to guidelines, significant gaps remain in ensuring human understanding of the solutions, particularly highlighted by discrepancies in proofs.
Key evidence
- OpenAI consulted an advisory group of nine prominent mathematicians but still fell short of their standards for human understanding in mathematical results.
- Only 42% of the proofs released by OpenAI underwent formalization, which the advisory group recommended for proofs that are not easily understood.
- A recent paper documented discrepancies between OpenAI's natural language proofs and the Lean code, raising concerns about the reliability of AI-generated solutions without human oversight.
Why it matters
The failure to meet community standards raises questions about the reliability of AI-generated mathematical solutions. Without human oversight, there is a risk that these solutions may not be accurately understood or applied, potentially hindering progress in the field. The advisory group's call for more formalization and human involvement underscores the importance of collaboration between AI and human mathematicians.
📖 Reader Mode
~4 min readWhen OpenAI released hundreds of claimed solutions to some of the world’s hardest math problems this week, the frontier lab said that it had consulted an advisory group of elite mathematicians to avoid the controversy that came with the last time one of its models solved a long-standing problem in the field.
But OpenAI fell short of those standards, particularly where the mathematicians emphasized the need for human understanding of a mathematical result. That’s especially concerning after a new paper highlighted gaps between the natural language and formally expressed solution to a million-dollar problem ostensibly solved by OpenAI’s models.
The Advisory Group on Mathematics and Artificial Intelligence, hosted by Princeton University’s Institute for Advanced Studies, is made up of nine prominent researchers at institutions around the world.
The organization released guidelines for frontier labs solving math problems at the end of September. In a statement on the latest set of proofs, the AGMAI said that “it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully.”
However, the organization’s first request was “to stop testing advanced mathematical problems on proprietary models.” OpenAI’s release explicitly says that it is evaluating its proprietary models using open research problems in mathematics.
The advisory group did not respond when asked by TechCrunch for a more thorough evaluation of OpenAI’s latest proof release. The lab clearly followed some of its principles, including releasing results as soon as possible and including information about how the models reached their conclusions. But not for all of them: Just ten of the 719 manuscripts included releases of the model’s chain of thought.
For papers that people don’t understand, the mathematicians suggested the proofs should be formalized — but just 42% of the proofs released by OpenAI had not undergone this process.
Ultimately, it’s still not clear that OpenAI is taking “responsibility for ensuring that human understanding will follow” when releasing its proofs, in accordance to the AGMAI principles. AGMAI suggested that OpenAI should help fund the work of human mathematicians who will be required to make the lab’s solutions meaningful in any real way.
“Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is ‘solved’, and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field,” Terence Tao, a prominent mathematician who has criticized OpenAI’s approach, wrote on social media after the release.
That problem is exemplified by a paper released this week by mathematicians at the University of Cambridge and King’s College in London that questions the way frontier labs are approaching these challenges.
When AI models solve mathematical problems, they first create a “natural language” explanation, then try to express that result in Lean, a programming language that in theory confirms the accuracy of the proof by compiling it as code.
However, there may be problems with the way the models translate their natural language proofs into code; this paper documents at least two discrepancies between the natural language proof and the Lean code behind the solution OpenAI has offered to a problem derived from the Navier-Stokes equations that describe the complex behavior of fluids.
These discrepancies don’t necessarily disprove either solution, but they do raise questions on whether we can simply rely on models to formalize their own solutions without human involvement. That’s one reason that AGMAI asked OpenAI to “include machine-readable metadata correlating the natural language and formal artifacts,” something that the frontier lab did not do with these releases.
“Because of the phenomenon of mistranslations — as highlighted in this paper — the NL proof by OpenAI and
other autoformalised Lean proofs should not prima facie be trusted without the same peer review process and
scrutiny that other proofs are subjected to,” the authors of the “lost in translation” paper conclude.
Mathematicians stress that when new results are discovered by humans, they take responsibility for them and engage with the broader community through papers, talks, and seminars. That process increases understanding of the solutions, finds strategies that can be used to solve other problems, and allows the new knowledge to be applied in practical fields.
When a model is prompted to solve a hard problem and spits out a solution, “there is not human understanding of them at the point of release, and now the work begins,” Harvard University mathematics professor Melanie Wood told TechCrunch.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Tim Fernholz is a journalist who writes about technology, finance and public policy. He has closely covered the rise of the private space industry and is the author of Rocket Billionaires: Elon Musk, Jeff Bezos and the New Space Race. Formerly, he was a senior reporter at Quartz, the global business news site, for more than a decade, and began his career as a political reporter in Washington, D.C. You can contact or verify outreach from Tim by emailing [email protected] or via an encrypted message to tim_fernholz.21 on Signal.
— Originally published at techcrunch.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from TechCrunch
See more →
AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors
AI chip startup Etched has achieved a $10.3 billion valuation after a $300 million Series C funding round, led by Sequoia and supported by notable investors like Andreessen Horowitz. The company claims to have developed innovative low-voltage chips for AI inference, significantly enhancing performance and reducing costs, with $1 billion in orders already booked.

