Transferable Guidance with Constrained Allocation for On-Policy Self-Distillation
Abstract
On-policy self-distillation uses training-only privileged information to supervise a model on its own responses. Its effectiveness depends on both the supervision targets and the allocation of training weight across response positions. We propose TransCAP to coordinate these choices through a shared task-conditioned residual. The residual contrasts task-plus-privilege predictions with a privilege-only reference, reweighting a task-only anchor and defining initial position priorities. A constrained KL projection then adjusts these priorities using local directional conflict at the current student distribution. The projection bounds measured directional conflict and prevents the allocation from becoming more concentrated than its initial form. On Qwen3.5-9B, TransCAP achieves reasoning and code averages of 80.65% and 72.22%, leading the strongest baselines by 6.58 and 4.32 percentage points. It also achieves state-of-the-art performance on Qwen3-4B and Qwen3.5-27B across both task groups.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.