
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Quick Answer
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours.
Quick Take
Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.
Key Points
- Claude Opus 4.7 achieved a 56% solve rate on the MirrorCode benchmark.
- The model rebuilt a 16,000-line toolkit in just 14 hours.
- All tested models failed on the most complex programming tasks.
- The single MirrorCode task cost $2,600 to run over 19 days.
- The results raise concerns about AI's cost-effectiveness in complex programming.
DeepSignal Analysis
What happened
Epoch AI's MirrorCode benchmark evaluated several AI models on programming tasks, revealing Claude Opus 4.7 as the top performer with a 56% solve rate. The model completed a complex bioinformatics toolkit task in 14 hours, costing $251. However, all models struggled with larger tasks, indicating ongoing limitations in AI programming capabilities.
Key evidence
- Claude Opus 4.7 achieved a 56% solve rate in the MirrorCode benchmark, outperforming GPT-5.5 and Gemini 3.1 Pro Preview.
- The AI model completed a 16,000-line bioinformatics toolkit task in 14 hours, while human engineers would require 2 to 17 weeks.
- The largest tasks in the MirrorCode benchmark remain unsolved by any of the tested models, highlighting significant challenges in AI programming.
Why it matters
The results from the MirrorCode benchmark raise questions about the cost-effectiveness of AI in software development, especially given the $2,600 expenditure for a single task run over 19 days. While progress is evident, the inability of models to tackle larger tasks suggests that AI is not yet a reliable substitute for human engineers in complex programming scenarios. This could impact investment and development strategies in the AI industry.
📖 Reader Mode
~2 min readThe 25 target programs cover Unix utilities, data serialization, bioinformatics, interpreters, static analysis, cryptography, and compression. Each AI-generated solution must exactly reproduce the output of the original program, including hidden end-to-end tests the model never sees during development.
Another difference from many other benchmarks is the inference budget. Existing software engineering benchmarks often cap costs at $1 to $10 per task, even when a human would need weeks to finish the same work, the developers write.
According to Epoch AI, one of the largest tasks in MirrorCode cost $2,600 for a single run. The AI worked continuously for 19 days with no human involvement at all.
Claude Opus 4.7 rebuilds a bioinformatics toolkit in 14 hours
Epoch AI says AI can already handle demanding long-term programming tasks. The standout example comes from Claude Opus 4.7, which reimplemented gotree, a bioinformatics toolkit with roughly 16,000 lines of Go code and over 40 commands. A human engineer working without AI help would need 2 to 17 weeks for the same job, according to the researchers. Opus 4.7 finished in 14 hours for $251.

In the overall rankings, Claude Opus 4.7 hit a solve rate of 56 percent. GPT-5.5 followed at 44 percent, and Gemini 3.1 Pro Preview came in at 32 percent. Even when models fail to fully reimplement a program, they typically pass 90 percent or more of the tests.
The hardest tasks still stump every model
Despite the progress, MirrorCode is far from solved. Tasks fall into three categories: small, medium, and large. Small programs like uuid or parseqsv get reliably reimplemented by all tested models. The largest tasks beat every model tested.

The researchers are still seeing rapid gains. Leading models from a year ago would have scored only about 30 percent and been limited to simpler programs like a calendar utility, Epoch AI says.
Cost trends don't follow a clear pattern. GPT-5.5 costs three times as much as GPT-5 for the same tasks, while Claude Opus 4.7 runs three times cheaper than Claude Opus 4.1.
Epoch AI has open-sourced the scaffold and 22 of the 25 target programs, covering 132 task instances across six programming languages. Three programs are kept private for testing.
The researchers point to one important caveat: since MirrorCode uses open-source programs as targets, the models may have already seen the original code during training. Initial tests suggest "the results were not dominated by memorization, but we cannot rule out the possibility that memorization contributes to AI performance," they write.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
Moonshot AI releases Kimi K3 open weights and infrastructure after shaking up the frontier model race
Moonshot AI has released Kimi K3's open weights and infrastructure, achieving 2.5 times more intelligence per compute unit. While it competes closely with models like Fable 5 and GPT-5.6 Sol on benchmarks, independent tests reveal significant gaps in cyber capabilities and math skills, suggesting reliance on distillation techniques.

