
Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it
Quick Answer
Andrew Ho, a former OpenAI researcher, predicts over $100 billion will be needed for targeted training data to improve large language models (LLMs) like GPT-5.6, which currently achieve only 30% success in complex bioinformatics tasks.
Quick Take
He argues that scaling alone won't suffice, as many economically relevant skills are underrepresented in existing datasets, leading to inconsistent performance in .
Key Points
- Ho claims large language models generalize poorly due to insufficient training data.
- Current models like GPT-5.6 only achieve about 30% success in bioinformatics tasks.
- Ho's products will focus on datasets for scientific analyses and everyday lab work.
- Cambridge researcher Adam Hunt notes LLMs are becoming more specialized, not versatile.
- Debate continues on whether LLMs can develop capabilities beyond their training data.
DeepSignal Analysis
What happened
Andrew Ho, a former OpenAI researcher, predicts that over $100 billion will be needed for targeted training data to enhance large language models (LLMs) like GPT-5.6, which currently show only 30% success in complex bioinformatics tasks. He argues that scaling alone is insufficient due to the underrepresentation of economically relevant skills in existing datasets.
Key evidence
- Ho claims that most economically relevant skills are poorly represented in current datasets, leading to inconsistent performance in LLMs.
- Current models like GPT-5.6 achieve only about a 30% success rate in complex bioinformatics tasks, according to Ho.
- Cambridge researcher Adam Hunt notes that while programming capabilities in LLMs are improving, areas like language quality and simple logic are stagnating or declining.
Why it matters
Ho's assertion highlights a significant challenge in the AI industry: the need for high-quality, targeted training data to improve model performance. As LLMs become more specialized, the economic viability of AI labs may be threatened if they cannot adapt to these data requirements. This situation raises questions about the sustainability of current AI development strategies and the long-term effectiveness of scaling models without addressing data quality.
Source Excerpt
Former OpenAI employee Andrew Ho and Cambridge researcher Adam Hunt see a growing problem with . Instead of becoming more versatile, the models are becoming more specialized, excelling at coding and math while stagnating or even regressing in other areas. Ho is leaving OpenAI to start a company focused on specialized training data and predicts that AI labs will need to spend more than $100 billion on targeted data collection.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

