
METR introduces a new metric to calculate exactly when AI agents become more expensive than humans
Quick Answer
METR introduces the 'expenditure horizon' metric to determine when AI agents become costlier than humans, revealing that humans spend about $2,500 for each 1% speedup.
Quick Take
In tests, AI models like GPT-5.5 and Opus-4.8 showed limited improvements, with expenditure horizons between $0 and $3,300, highlighting the need for future models to enhance efficiency.
Key Points
- METR's expenditure horizon indicates AI is cheaper below $2,500 per 1% speedup.
- AI models GPT-5.5 and Opus-4.8 achieved minimal improvements in tests.
- Humans spent about 16 hours for each 1% speedup at $150/hour.
- Recent models like Opus 5 may significantly shift the expenditure horizon.
- The study highlights the limitations of AI working independently without human collaboration.
DeepSignal Analysis
What happened
METR has introduced a metric called the 'expenditure horizon' to evaluate when AI agents become more expensive than human labor. In their tests, humans require approximately $2,500 for each 1% speedup, while AI models like GPT-5.5 and Opus-4.8 showed limited improvements, with expenditure horizons ranging from $0 to $3,300.
Key evidence
- METR's expenditure horizon metric indicates the cost point where AI and human labor yield the same improvement, with humans spending about $2,500 for each 1% speedup.
- In tests involving six AI models, only GPT-5.5 and Opus-4.8 achieved meaningful expenditure horizons, while others like GPT-5 and Opus-4.1 showed no real progress.
- The study highlights that METR's findings are based on older AI models, and newer models like Opus 5, which reportedly perform better, were not included in the analysis.
Why it matters
This metric provides a nuanced understanding of the cost-effectiveness of AI versus human labor, revealing that while AI can excel in simple tasks, it struggles with more complex challenges. The findings underscore the importance of improving AI efficiency to make it a viable alternative to human labor, especially as tasks become more demanding. Additionally, the study emphasizes the need for further research into hybrid human-AI collaboration, which could potentially yield better results than either working independently.
Source Excerpt
METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

