
New benchmark exposes how badly AI struggles with real knowledge work
Quick Answer
A new benchmark reveals that even top AI models, like those from leading companies, only solve 3% of realistic knowledge work tasks.
Quick Take
This stark performance gap highlights the limitations of current AI technologies in practical applications, affecting industries reliant on knowledge work.
Key Points
- Top AI models struggle with realistic knowledge work, solving only 3% of tasks.
- The benchmark highlights significant limitations in AI's practical applications.
- Industries relying on knowledge work may face challenges due to AI performance.
- Current AI technologies are not yet equipped for complex knowledge tasks.
📖 Reader Mode
~1 min readEven the best AI model fails at realistic knowledge work, fully solving just 3 percent of tasks.
The new AA-Briefcase benchmark from Artificial Analysis puts AI models through multi-week knowledge work projects built from thousands of fragmented source files like Slack threads, emails, meeting transcripts, and large data exports. The top performer, Claude Fable 5, hits the highest rubric pass rate but still nails all criteria on just 3 percent of tasks. On 31 out of 91 tasks, no model even clears 50 percent.

The types of errors shift as models get better. Weaker models choke on basic execution as they miss relevant files or spit out unusable results. Stronger models fail more quietly, as they hit the obvious requirements but miss details you'd only catch by piecing together information from multiple sources.
There also is a significant price gap: Per-task costs span more than 800x, from about $0.04 for DeepSeek V4 Flash to over $31 for Claude Fable 5.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

