
Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it
Quick Answer
Andrew Ho, a former OpenAI researcher, predicts over $100 billion will be needed for targeted training data to improve large language models (LLMs) like GPT-5.6, which currently achieve only 30% success in complex bioinformatics tasks.
Quick Take
He argues that scaling alone won't suffice, as many economically relevant skills are underrepresented in existing datasets, leading to inconsistent performance in .
Key Points
- Ho claims large language models generalize poorly due to insufficient training data.
- Current models like GPT-5.6 only achieve about 30% success in bioinformatics tasks.
- Ho's products will focus on datasets for scientific analyses and everyday lab work.
- Cambridge researcher Adam Hunt notes LLMs are becoming more specialized, not versatile.
- Debate continues on whether LLMs can develop capabilities beyond their training data.
DeepSignal Analysis
What happened
Andrew Ho, a former OpenAI researcher, predicts that over $100 billion will be needed for targeted training data to enhance large language models (LLMs) like GPT-5.6, which currently show only 30% success in complex bioinformatics tasks. He argues that scaling alone is insufficient due to the underrepresentation of economically relevant skills in existing datasets.
Key evidence
- Ho claims that most economically relevant skills are poorly represented in current datasets, leading to inconsistent performance in LLMs.
- Current models like GPT-5.6 achieve only about a 30% success rate in complex bioinformatics tasks, according to Ho.
- Cambridge researcher Adam Hunt notes that while programming capabilities in LLMs are improving, areas like language quality and simple logic are stagnating or declining.
Why it matters
Ho's assertion highlights a significant challenge in the AI industry: the need for high-quality, targeted training data to improve model performance. As LLMs become more specialized, the economic viability of AI labs may be threatened if they cannot adapt to these data requirements. This situation raises questions about the sustainability of current AI development strategies and the long-term effectiveness of scaling models without addressing data quality.
📖 Reader Mode
~3 min readAndrew Ho is leaving OpenAI after just eight months, convinced that large language models generalize poorly. Even in well-funded areas like programming, he says, their performance is inconsistent.
The root cause, according to Ho, is a lack of training data. Most skills that matter economically are barely represented in existing datasets.
"Most work is highly contextual and not easily encoded into a gradable environment; even if we can observe a 'golden path' taken by a human which we believe to be good, it's challenging to understand whether alternate, counterfactual paths produce good or bad outcomes," Ho writes.
Scaling alone won't fix this, he argues, and expects AI labs will have to spend more than $100 billion on targeted data collection in the years ahead. Ho is also skeptical of the sky-high valuations at frontier labs like OpenAI or Anthropic, which he says are chronically unprofitable because they have to keep pouring growing sums into new models just to stay ahead of cheaper rivals like Qwen or Kimi.
His first products target two areas. One is datasets for complex scientific analyses in bioinformatics, where even current models like GPT-5.6 Sol hit only about a 30 percent success rate, a topic he worked on at OpenAI. The other is datasets for everyday lab work, such as when researchers submit photos of experiments to AI models for evaluation. Chemistry, materials science, healthcare, and broader knowledge work are planned to follow.
Latest models may be getting sharper in some areas while dulling in others
Cambridge researcher Adam Hunt shares Ho's skepticism. Hunt describes how his own view of LLMs has shifted from optimistic to increasingly pessimistic. His argument is that the latest models aren't becoming more versatile but more specialized. Programming and complex math capabilities keep improving, while areas like language quality and simple logic are stagnating or even getting worse, which lines up with Ho's observation of uneven performance.

The reason, Hunt says, is that reinforcement learning works well in a domain like code because clear reward signals and complete training data exist there. In other areas, that data simply isn't available. The early progress of large language models and their apparent ability to generalize through sheer scale and reasoning was a side effect of training on a broad text corpus. It wasn't a sign of true general understanding.
The generalization question is far from settled
This debate keeps coming back to the same question: can language models develop capabilities that go beyond reproducing and recombining their training data, especially in areas where results can't be automatically verified? Scientists don't agree on the answer.
Hunt himself puts his confidence at only about 40 percent and acknowledges that technical advances could prove him wrong, for example if specialized AI models can be combined into something that resembles more general intelligence.
In a recent position paper titled "LLMs can't jump," Google Deepmind's Tom Zahavy offers a structural explanation for why models are uneven. Language models are good at deduction and induction but fail at creative abduction, meaning the ability to invent a cause for which no linguistic precedent exists yet. As a possible fix, Zahavy points to action-controllable world models that allow for counterfactual experiments. LLMs could still play a key role in such advanced systems, even if they hit their limits when working alone.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

