acceptodds
Under review as a conference paper at ICLR 2027

SPRA: Scaffold-Preserving Representation Alignment for Robust Socratic Tutors

Abstract

Large language model (LLM)-based Socratic tutors are increasingly employed to guide student learning through multi-turn questioning, yet they remain vulnerable to scaffolding collapse, where sustained student pressure causes the tutor to abandon Socratic guidance and reveal direct answers. Existing defenses mainly regulate surface outputs through prompting, preference tuning, or filtering, but they do not address the hidden-state drift associated with such trajectory-level failures. Motivated by empirical evidence that scaffolding collapse is associated with systematic hidden-state drifts, we propose a Scaffold-Preserving Representation Alignment (SPRA) framework that integrates supervised fine-tuning, trajectory-weighted direct preference optimization (DPO), and a margin-preserving representation loss that preserves the hidden-state projection margin between scaffolding and collapse responses. We evaluate SPRA across five STEM disciplines under a multi-turn red-teaming protocol. On Qwen3-8B, SPRA keeps the Collapse Rate at or below 32% in every discipline, delays the average collapse onset beyond 9 turns, and maintains a low over-refusal rate, suggesting that representation-level alignment can improve the scaffolding robustness of Socratic tutors under our settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.