It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
Quick Answer
The W2SPO method enhances off-policy reinforcement learning by integrating short auxiliary segments into target model trajectories, achieving a Pass@1 improvement from 62.3% to 64.2% and a 3.55x training speedup on mathematical reasoning tasks with 4B scale models.
Key Points
- W2SPO uses 8-token auxiliary segments to inform policy exploration.
- Achieves superior performance on mathematical reasoning benchmarks with 4B scale models.
- Outperforms post-trained baselines in terms of policy updates.
- Improves training speed by 3.55 times compared to vanilla .
- Addresses semantic redundancy in target model samples during reasoning tasks.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self generated rollouts. However, we identify a critical support limited bottleneck in this paradigm: on challenging reasoning tasks, the target model's samples often exhibit semantic redundancy, converging into the same erroneous "reasoning basins" that offer negligible reward contrast for policy updates. In this paper, we propose to overcome this limitation through a weak to strong learning paradigm, where a policy's exploration is informed by a weaker but computationally efficient auxiliary model. We introduce W2SPO, an off policy RL method that injects short auxiliary segments often as brief as 8 tokens into intermediate target model trajectories and the target model then completes the reasoning path from these diverted states. Policy updates are restricted to these short inserted segments based on final verifiable rewards. Empirically, W2SPO achieves superior performance among evaluated 4B scale models on mathematical reasoning benchmarks, outperforming evaluated post trained baselines. Compared with vanilla GRPO under the same sampling budget, W2SPO improves Pass@1 from 62.3% to 64.2% while achieving a 3.55 times training speedup. These results suggest that weak auxiliary branches can induce stronger target reasoning policies by expanding local exploration support.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.16205 [cs.AI] |
| (or arXiv:2607.16205v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16205 arXiv-issued DOI via DataCite |
Submission history
From: Dayu Wang [view email]
[v1]
Thu, 7 May 2026 13:16:09 UTC (457 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.