
Some mathematicians call for OpenAI boycott after AI-generated proofs flood their field
Quick Answer
Mathematicians, led by Fields Medalist Terence Tao, call for a boycott of OpenAI after its AI-generated proofs, including claims of solving over 100 math problems, disrupt traditional research and collaboration, raising concerns about attribution and the future of mathematical inquiry.
Key Points
- OpenAI claims its AI model solved over 100 math problems in one month.
- The AHM criticizes OpenAI for ignoring advice on advanced math problem testing.
- Tao warns that AI-generated solutions could undermine traditional mathematical collaboration.
- Concerns arise over attribution and plagiarism in AI-generated proofs.
- The 'Math 2.0' era should prioritize explanation and community over mere problem-solving.
DeepSignal Analysis
What happened
A group of mathematicians, led by Terence Tao, has called for a boycott of OpenAI due to concerns over AI-generated proofs disrupting traditional mathematical research. OpenAI claimed its model solved over 100 math problems, raising issues of attribution and the future of mathematical inquiry.
Key evidence
- OpenAI claimed its internal model solved more than 100 open math problems in a single month, including an approach to the Navier-Stokes Millennium Problem.
- The AHM criticized OpenAI for releasing over 700 files at once, stating it was 'not a demonstration of scholarship, but a demonstration of power.'
- Terence Tao and 24 other Fields Medalists warned of a 'severe misalignment' between AI industry goals and the true objectives of mathematics.
Why it matters
The mathematicians' call for a boycott highlights a growing concern that AI-generated solutions could undermine the collaborative and conceptual nature of mathematical research. The rapid pace of AI advancements may lead to a disconnect between traditional mathematical inquiry and the methods employed by AI, potentially stifling innovation and understanding in the field.
What to watch
📖 Reader Mode
~6 min readFields Medalist Terence Tao, who chairs the group, shared the statement as a guest post on his blog, while also outlining a "Math 2.0" era on Mastodon in which solving problems should no longer be the main focus of the discipline. Complexity theorist Scott Aaronson described the situation as a "Mathocalypse" on his blog.
The AHM statement escalates a months-long debate. OpenAI had already caused a stir by claiming an internal model solved more than 100 open math problems in a single month, including an approach to the Navier-Stokes Millennium Problem. In response to growing criticism, OpenAI set up the advisory group AGMAI at the Institute for Advanced Study, though the group explicitly has no say over the pace of the company's internal research.
AHM sees the release as a show of force
The AHM says OpenAI has already ignored the central premise of AGMAI's advice, which held that advanced math problems shouldn't be tested on internal models. "Mathematicians did not ask for this work to be done," the statement reads, describing the release of more than 700 files at once as "not a demonstration of scholarship, but a demonstration of power."
The statement's opening paragraph also brings up the many copyright lawsuits OpenAI is fighting around the world, suggesting without saying so directly that the company could only achieve these results because it trained on mathematicians' work, whether legally or not.
That same dispute surfaced when OpenAI recently announced a solution to the Navier-Stokes problem shortly before two mathematicians could present their own AI-assisted solution. The two had used ChatGPT in their work, and OpenAI denied suspicions that it had used data from those interactions to train its own system and reach a solution faster.
Aaronson contrasts this with what he calls the "Anthropic model." OpenAI releases raw AI proof drafts in one batch, triggering a race to work through them. Anthropic instead partnered with two algorithm researchers whose AI model supplied the key idea for disproving two decades-old conjectures. The researchers received compensation and wrote a version of the proof that humans could follow.
Both approaches have drawbacks. OpenAI's method leaves the community doing the thankless work of making proofs readable, for free. Anthropic's method lets a private company choose which mathematicians get to serve as "emissaries" for a result.
AI is cracking open math problems faster than anyone can verify the results
Tao previously joined 24 other Fields Medalists, including Peter Scholze, Maryna Viazovska, and Martin Hairer, in signing a statement warning of a "severe misalignment" between the AI industry's goals and those of mathematics. Problem-solving, they argue, is only a tool for the discipline's real goal of conceptual understanding, and mass-producing solved problems at an ever-faster pace could "destroy fertile ground instead of breathing life into new ideas."
The signatories also criticize rushed announcements of AI-generated solutions without proper write-ups or citations of relevant prior work, saying this raises "severe attribution and plagiarism questions."
On Mastodon, Tao now goes further. In traditional mathematics, proofs of long-standing open problems would lead to talks, workshops, new collaborations, and inclusion in textbooks, attracting young researchers who would work to understand and contextualize them.
Now he sees the opposite. Problems are being solved autonomously by AI users who have no interest in the field and can't understand the results well enough to answer questions or give talks, producing far fewer seminars and collaborations than traditional breakthroughs would. Researchers are also holding back promising directions for fear that competitors will scoop the solution.
For Tao, the damage is irreversible. Once a problem is considered solved, it can't be made unsolved again. Even knowing a solution exists can "contaminate," as he puts it, the search for other approaches that might yield further insights. Solutions are being "harvested" on a large scale, leaving entire branches of mathematics less fertile than before.
Tao argues that the "Math 2.0" era must stop treating problem-solving as its main measure of progress and start valuing explanation, community-building, and the opening of new research directions, with corresponding changes to how training, publication, and career advancement are judged.
This builds on Tao's earlier vision of "industrial mathematics" from 2024, though humans still set the pace in that version. AI was supposed to assist mathematicians much as engines assist chess players, enabling broad, relatively superficial research that complemented the deeper work of human experts. Today, Tao sees those roles reversed. AI is solving the deep problems while humans work through the results afterward. Like many other experts, Tao apparently badly underestimated the pace of AI development.

Three hours of AI compute can now rival a mathematician's entire career
Aaronson shows what this means for individual researchers through the experience of his wife, Dana Moshkovitz, who has devoted her entire career to the Unique Games Conjecture, an open problem in theoretical computer science. The release includes a claimed proof of it.
In text messages to Aaronson on the night of the release, she wrote, "It feels like something written by someone who's on psychedelics." Much of the document was unclear. The paper cited numerous works without explaining why they applied, even though earlier results should have ruled that out.
"Basically the paper is so horribly written that it's impossible to read it without AI help," Moshkovitz said. The proof relied on an entirely new, bizarre construction that she described as "some alien craziness."
Aaronson lists other results from complexity theory, number theory, and algorithm research, including partial progress on several remaining Clay Millennium Problems. Any one of them, in his view, would have ranked among the year's top results on its own.
One area is notably absent from the release, however. Cryptography. Citing unnamed sources, Aaronson reports that AI companies are now quietly probing weaknesses in cryptographic protocols and primitives. About 8,000 problems were tested overall, with a success rate of roughly five percent and an average of three hours of compute at GPT-Pro level per problem. The model used was likely OpenAI's current internal one as well, which could ship to paying ChatGPT customers in the coming months.
Mathematicians disagree over whether a boycott will help
Comments on Tao's blog post show how divided the field is. Several commenters back the boycott and propose that researchers stop using OpenAI products while also calling for better protection of preprint servers like arXiv against mass collection of training data.
Others consider the AHM's position unrealistic, arguing that OpenAI won't stop working on mathematics and that public results are better than private ones. Another commenter asks whether a problem can really count as solved when neither the authors nor anyone else fully understands the proof. In the case of Navier-Stokes, it remains unclear whether the result actually deepens anyone's understanding of fluid dynamics.
AGMAI itself takes a more diplomatic stance than the AHM, viewing the release as a first step with the mathematical understanding of the results only now beginning. Still, the group warns that the future of math research can't consist of working through results produced by AI labs.
"Mathematicians must be able to formulate their own questions, develop their own approaches, and explore directions that have not been selected as examples of an AI system's capabilities. Equitable access to powerful research tools and adequate computational resources are essential to that freedom," the organization writes.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

