CurveShift: Is Agent Progress Scalar? Separating Level from Shape
Quick Answer
The study reveals that progress in large language models, particularly post-September 2024, shows a significant increase in solving hard problems, with a +0.40 logits improvement, raising the solve rate from 18% to 25%.
Quick Take
This gain is attributed to stronger reasoning models and is isolated using the benchmark, which avoids confounding factors from agentic harnesses.
Key Points
- Models released after September 2024 show a +0.40 logits gain on hard tasks.
- The hard-problem solve rate increased from approximately 18% to 25%.
- LiveCodeBench benchmark isolates model performance without agentic scaffolding.
- The effect is primarily driven by advanced reasoning models.
- Results are specific to competitive programming tasks.
DeepSignal Analysis
What happened
The study investigates the performance of large language models, particularly after September 2024, revealing a +0.40 logits improvement in solving hard problems. This increase raises the solve rate from 18% to 25%, attributed to stronger reasoning models. The findings are based on the LiveCodeBench benchmark, which isolates the effects of agentic harnesses.
Key evidence
- Models released after September 2024 show a +0.40 logits improvement in solving hard problems, increasing the solve rate from 18% to 25%.
- The LiveCodeBench benchmark was used to isolate model performance from confounding factors related to agentic harnesses.
- The study indicates that most apparent gains in hard tasks do not reflect a change in the difficulty-response curve but are largely due to ceiling effects.
Why it matters
Understanding the true capabilities of large language models is crucial for their application in complex problem-solving. The findings challenge the notion that improvements in performance are uniformly distributed across task difficulties. By isolating the effects of different models, the study provides clearer insights into the actual advancements in reasoning capabilities, which can inform future research and development in AI.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Hanwen Xing, Pengyun Wang, BingXu Meng, Kumail Alhamoud, Xiang Li, Jicheng Wang, Xin Yu, Xinyang Han, Xiaomin Li, Philip Torr, Yuexing Hao
Abstract:Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that most of the apparent shift in gains toward harder tasks does not reflect a change in the shape of the difficulty-response curve. On METR time-horizon data, a single Rasch model with rising ability reproduces this pattern, so it is largely explained by ceiling effects rather than a qualitative change in capability. This echoes how the choice of metric can make claimed emergent abilities look like a property of the models themselves. We then identify a smaller hard-task effect that survives this control. Isolating it is difficult on agentic benchmarks, because newer models are usually run with newer agentic harnesses, so a gain on hard tasks cannot be assigned to the model or its scaffold. We break the confound with LiveCodeBench, a public competitive programming benchmark that runs no agentic scaffold while pairing dated models with an exogenous difficulty ordering. After accounting for the rise in overall ability, models released after September 2024 still gain on the hardest problems beyond what their easy and medium performance predicts, by about +0.40 logits under our most conservative assumption, raising the hard-problem solve rate from roughly 18% to 25%. The effect is led by the strongest reasoning models and holds for hard tasks that need only short reasoning, not autonomy over long horizons. We present this as a result specific to competitive programming, since our clean identification rests on a single coding benchmark. We release the LiveCodeBench Difficulty Panel (66 dated models x 1,055 problems) and our analysis code.
| Comments: | 25 pages, 4 figures, 7 tables. Data and code: this https URL |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| ACM classes: | I.2.7; I.2.6; G.3 |
| Cite as: | arXiv:2608.00355 [cs.CL] |
| (or arXiv:2608.00355v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.00355 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hanwen Xing [view email]
[v1]
Fri, 31 Jul 2026 23:58:25 UTC (128 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.