
Language models can't spark scientific revolutions, but world models might
Quick Answer
Zahavy argues that while language models excel in deduction and induction, they struggle with 'manipulative abduction,' the creative leap needed for scientific breakthroughs.
Quick Take
Models like AlphaProof and GPT-5 show promise in formal derivation, but lack the sensory grounding that enabled Einstein's insights. To overcome this, Zahavy suggests developing action-controllable world models that allow for counterfactual experimentation.
Key Points
- Language models excel at deduction and induction but struggle with creative reasoning.
- AlphaProof and GPT-5 achieve high scores on mathematical problems but lack sensory grounding.
- Zahavy highlights Einstein's insights as examples of 'manipulative abduction' in science.
- Action-controllable world models could enable new scientific axioms through counterfactual experiments.
- Existing AI systems like AlphaEvolve optimize but cannot create entirely new frameworks.
DeepSignal Analysis
What happened
Zahavy argues that while language models excel at deduction and induction, they struggle with manipulative abduction, which is crucial for scientific breakthroughs. He highlights that models like AlphaProof and GPT-5 can achieve high scores in formal derivation but lack the sensory grounding necessary for creative leaps. To address this, he suggests developing action-controllable world models that enable counterfactual experimentation.
Key evidence
- Zahavy identifies three types of reasoning: deduction, induction, and abduction, with machines currently excelling in the first two but struggling with manipulative abduction.
- Models like AlphaProof, Gemini, and GPT-5 have achieved gold-level scores on International Mathematical Olympiad problems, demonstrating their capabilities in formal derivation.
- Zahavy proposes action-controllable world models, such as Genie, which allow for counterfactual experiments, potentially enabling the creative leaps necessary for scientific innovation.
Why it matters
The distinction between reasoning types is significant for understanding the limitations of current AI models in scientific discovery. While language models can process and derive existing knowledge, their inability to perform manipulative abduction indicates a gap that could hinder future advancements in AI-driven science. Developing world models that allow for experimentation could bridge this gap, fostering innovation.
Source Excerpt
Can language models spark a scientific revolution? In a position paper titled " can't jump," Google Deepmind's Tom Zahavy argues they can't. They're missing the cognitive mechanism needed to create something truly new.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

