The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation
Quick Answer
The study evaluates Qwen 3.6 Plus and MiniMax M2.5 across three harnesses, revealing that harness choice can induce a 40x difference in tokens per task solved, while pass-rate differences remain minimal.
Quick Take
This highlights the importance of considering harness-model pairs for accurate coding agent evaluations, as they significantly affect real-world costs and performance metrics.
Key Points
- Harness choice can lead to a 40x difference in tokens per solved task.
- Pass-rate differences between models remain between 0-8 percentage points.
- Failure fingerprints indicate harness-level biases that are model-independent.
- Selecting harness-model pairs by pass rate is recommended for better evaluations.
- The study provides anonymized configs and raw trial logs for further analysis.
DeepSignal Analysis
What happened
A study evaluated the performance of coding agents Qwen 3.6 Plus and MiniMax M2.5 across three different harnesses. The findings indicated that the choice of harness could lead to a significant difference in tokens required per task solved, with variations up to 40 times, while pass rates showed minimal differences.
Key evidence
- The study assessed Qwen 3.6 Plus and MiniMax M2.5 using three open-source harnesses: Goose, OpenCode, and OpenHands-SDK.
- Harness choice resulted in up to a 40x difference in tokens per solved task, while pass-rate differences were only 0-8 percentage points.
- Failure patterns were consistent across models, indicating biases at the harness level that were largely independent of the models used.
Why it matters
This research emphasizes the need for a more nuanced evaluation of coding agents, as traditional metrics like model name and pass rate do not capture the full picture. The significant impact of harness choice on performance metrics suggests that evaluations could misrepresent a model's capabilities if harness effects are not considered. This has implications for real-world applications, where costs and performance can vary dramatically based on the selected harness.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified. Model-to-model comparison is valid when the harness is fixed; when it varies, performance and efficiency conflate model and scaffold effects. We evaluate Qwen 3.6 Plus and MiniMax M2.5 across three open-source harnesses (Goose, OpenCode, OpenHands-SDK) on a stratified 50-task subset of Terminal-Bench Pro. Harness choice induces up to a 40x difference in tokens per solved task, while paired within-model pass-rate differences remain 0-8 percentage points (95% paired-task bootstrap CIs include zero except for the largest gap). Failure fingerprints replicate across models (REASON for Goose, VERIFY/MAX_TURNS for OpenHands-SDK, idle-loop/TIME for OpenCode), indicating harness-level biases that are largely model-independent. For human-centered coding-agent evaluation, model name alone is an incomplete comparison unit: harness-model pairs determine real-world cost, latency, and oversight burden; no-action turns are a per-task wait tax, not just a token tax. We therefore recommend selecting harness-model pairs by pass rate under token/latency budgets, and reporting token usage, latency, and full harness specifications alongside any model comparison. We release anonymized configs, raw trial logs, aggregated snapshots, and analysis scripts.
| Comments: | 7 pages, 2 figures, 6 tables. Preliminary work; under review at the 5th DL4C Workshop @ ICML 2026 |
| Subjects: | Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.2; D.2.5 |
| Cite as: | arXiv:2607.22585 [cs.AI] |
| (or arXiv:2607.22585v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22585 arXiv-issued DOI via DataCite |
Submission history
From: Naman Vats [view email]
[v1]
Mon, 8 Jun 2026 20:27:06 UTC (37 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.