Stochastic Teacher Intervention for Agentic On-Policy Distillation
Quick Answer
The paper introduces STI-OPD, a stochastic teacher intervention framework that enhances on-policy distillation (OPD) for multi-turn agentic tasks.
Quick Take
By using teacher-student policy discrepancies to guide interventions, STI-OPD significantly improves performance on benchmarks for tool-integrated reasoning and long-horizon interactions, outperforming previous OPD methods across various student sizes.
Key Points
- STI-OPD uses teacher intervention to correct student actions in multi-turn tasks.
- The framework employs KL divergence to estimate policy discrepancies for adaptive interventions.
- Importance-Weighted Reverse KL objective preserves OPD goals amidst mixed-policy trajectories.
- STI-OPD outperforms previous OPD baselines across all evaluated benchmarks and student sizes.
- The approach addresses error accumulation in agentic decision-making over multiple turns.
Paper Resources
📖 Reader Mode
~2 min readAbstract:On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
| Comments: | Work in progress |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10878 [cs.CL] |
| (or arXiv:2610.10878v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10878 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Junnan Liu [view email]
[v1]
Wed, 7 Oct 2026 20:29:29 UTC (589 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
MARS-Gov introduces a framework for detecting bureaucratic bias in Dutch government documents, achieving a new state-of-the-art F1 score of 0.880. This model outperforms existing zero-shot detectors by 20.2 points and reduces unnecessary interventions to just 2.5%. The framework's dynamic '10th juror' adapts to emerging biases, enhancing legal language processing.