
AI agent teams waste massive tokens for barely measurable quality gains, research finds
Quick Answer
Research by Vals AI reveals that AI agent teams, including GPT-6 Sol and Claude Opus 5.5, yield minimal quality improvements while incurring costs 1.8x to 5.1x higher than single agents.
Quick Take
Only one out of four comparisons showed a significant advantage, suggesting that the expense of is often unjustified, especially when models are already optimized.
Key Points
- GPT-6 Sol's team scored 7.3 points higher at medium reasoning, but no advantage at maximum.
- Agent teams cost significantly more but provide negligible score improvements.
- OpenAI's Noam Brown confirms multi-agent systems mainly enhance speed, not quality.
- Coordination issues in large agent swarms lead to diminishing returns, termed 'coordination tax.'
- Fable 5.1 showed some gains but still underperformed compared to Opus 5.5 in tests.
DeepSignal Analysis
What happened
Research from Vals AI indicates that AI agent teams, such as GPT-6 Sol and Claude Opus 5.5, provide minimal quality improvements compared to single agents while incurring significantly higher costs. Only one out of four comparisons showed a notable advantage for teams, raising questions about the efficiency of multi-agent systems.
Key evidence
- Vals AI's tests revealed that teams cost between 1.8x and 5.1x more than single agents, with only one significant improvement noted.
- In a specific comparison, GPT-6 Sol's team scored 7.3 points higher at medium reasoning, but no advantage was observed at maximum reasoning.
- OpenAI's Noam Brown stated that multi-agent systems primarily enhance speed rather than quality, with four agents solving tasks twice as fast but also costing twice as much.
Why it matters
The findings challenge the justification for deploying multi-agent systems, especially when existing models are already optimized. The high costs associated with these systems may not translate into meaningful performance gains, suggesting a need for reevaluation in AI deployment strategies.
What to watch
📖 Reader Mode
~2 min read
AI agent teams deliver almost no better results than single agents, research finds.
Evals company Vals AI tested GPT-6 Sol and Claude Opus 5.5 on the "Vibe Code Bench," both solo and as teams, at two reasoning levels: medium and maximum reasoning effort. The teams cost between 1.8x and 5.1x more than single agents.
Out of four comparisons between teams and solo agents, only one showed a statistically significant improvement: GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher. At maximum reasoning, the team setup gave neither Sol nor Opus 5.5 any real advantage. The results suggest that the extra cost of agent teams isn't worth it in most cases, especially when models are already running at full compute.

Anthropic saw quality gains shrink as it added more agents in two of its own tests with Opus 5.5. Larger teams reached a given performance level faster, but going from ten to 100 agents only nudged scores up slightly after 24 hours. In separate ProgramBench tests, speed gains came with higher token usage.
| Task with Opus 5.5 | 1 agent | 10 agents | 30 agents | 100 agents |
|---|---|---|---|---|
| Knowledge base | 0.53 | 0.70 | 0.71 | 0.74 |
| Lean Theorem Proving | 0.39 | 0.66 | 0.66 | 0.68 |
Fable 5.1 showed stronger quality gains on the Lean theorem proving task above ten agents, but still scored below Opus 5.5 across all tests. On the knowledge base task, Fable's score actually dipped slightly when scaling from 30 to 100 agents.
More agents buy speed but not better output
OpenAI researcher Noam Brown confirmed in the Dwarkesh Podcast that multi-agent systems mainly buy speed, not better quality. Four agents solved tasks twice as fast but also cost twice as much. At 16 agents, the pattern held but grew slightly less efficient.
The effect depends heavily on the task. Web research and math parallelize well, but writing a novel doesn't, he said. Throwing 10,000 agents at a novel would be just as pointless as throwing 10,000 people at it. Brown acknowledged that scaling to very large numbers of agents remains largely unexplored because the costs are simply too high.
OpenAI developer Eric Provencher recently warned against using agent swarms for exactly this reason. They're most likely wasted money, he argued, because coordination between agents breaks down. He called it the coordination tax.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

