Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment
Quick Answer
The proposed Scaffold-Preserving Representation Alignment framework enhances Socratic tutoring by reducing the collapse rate to 32% on Qwen3-8B, delaying collapse onset beyond nine turns, and maintaining low over-refusal rates, thereby improving long-horizon tutoring robustness against red-teaming attacks.
Key Points
- Introduces a two-stage framework for Socratic tutors to prevent scaffolding collapse.
- Combines supervised fine-tuning with trajectory-weighted preference optimization.
- Achieves a 32% collapse rate and delays onset beyond nine dialogue turns.
- Evaluated across five STEM disciplines and various red-teaming strategies.
- Demonstrates improved robustness of long-horizon Socratic tutoring.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language model (LLM)-based Socratic tutors increasingly guide students through multi-turn questioning, but they can suffer from scaffolding collapse: under sustained student pressure, a tutor gradually abandons guided inquiry and reveals solutions directly. Prior defenses primarily constrain observable responses through prompting, preference optimization, or filtering, leaving the internal representation drift that precedes trajectory-level collapse largely unaddressed. We propose Scaffold-Preserving Representation Alignment, a two-stage framework that first warms up a Socratic tutor with supervised fine-tuning, then combines trajectory-weighted direct preference optimization with a margin-preserving representation loss anchored to frozen reference states. Our method is designed to maintain separation between scaffold-preserving and collapse-inducing hidden states across dialogue turns. We evaluate our method across five STEM disciplines and five red-teaming attack strategies. On Qwen3-8B, our method lowers Collapse Rate to 32%, delays average collapse onset beyond nine turns, and keeps over-refusal low, suggesting that representation-level alignment can improve the robustness of long-horizon Socratic tutoring under our red-teaming protocol.
| Comments: | preprint, under review |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.19371 [cs.AI] |
| (or arXiv:2607.19371v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.19371 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jing Shao [view email]
[v1]
Mon, 15 Jun 2026 07:55:09 UTC (760 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.