Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
Quick Answer
The DASH method improves reasoning in language models by segment-level credit assignment, reducing overthinking behaviors and achieving 50.8% accuracy on AIME25 benchmarks compared to 45.4% for GRPO.
Quick Take
This approach identifies productive self-reflection without costly annotations, enhancing performance in competitive math tasks.
Key Points
- DASH assigns credit based on reasoning segment contributions towards correctness.
- The method reduces unproductive self-reflection in language models.
- Achieved 50.8% accuracy on AIME25, outperforming by 5.4%.
- Intermediate answer commitments serve as a low-cost proxy for reflection evaluation.
- DASH enhances self-correction capabilities in reasoning tasks.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without improving answers. We show that these behaviors are not merely a consequence of length; even when controlling for response length, incorrect traces exhibit higher rates of unproductive self-reflection than correct ones. Addressing this requires identifying where self-reflection helps vs hurts, but obtaining these step-level annotations is costly. We observe that intermediate answer commitments within reasoning traces can provide a cheap proxy: by comparing each final answer candidate in the trace to the ground truth, we can determine whether subsequent reflection is productive without any additional supervision. Building on this insight, we propose DASH (Drift Aware advantage SHaping), which assigns segment-level credit based on whether each reasoning segment leads toward or away from correctness. On competition-level math benchmarks, DASH achieves the highest accuracy where overthinking is prevalent (AIME25: 50.8% vs. 45.4% GRPO) while reducing overthinking behaviors and achieving more productive self-correction than baselines.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.00482 [cs.CL] |
| (or arXiv:2607.00482v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.00482 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chia-Hsuan Lee [view email]
[v1]
Wed, 1 Jul 2026 06:09:56 UTC (1,279 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.