acceptodds
Under review as a conference paper at ICLR 2027

Transferable Guidance with Constrained Allocation for On-Policy Self-Distillation

Abstract

On-policy self-distillation uses training-only privileged information to supervise a model on its own responses. Its effectiveness depends on both the supervision targets and the allocation of training weight across response positions. We propose TransCAP to coordinate these choices through a shared task-conditioned residual. The residual contrasts task-plus-privilege predictions with a privilege-only reference, reweighting a task-only anchor and defining initial position priorities. A constrained KL projection then adjusts these priorities using local directional conflict at the current student distribution. The projection bounds measured directional conflict and prevents the allocation from becoming more concentrated than its initial form. On Qwen3.5-9B, TransCAP achieves reasoning and code averages of 80.65% and 72.22%, leading the strongest baselines by 6.58 and 4.32 percentage points. It also achieves state-of-the-art performance on Qwen3-4B and Qwen3.5-27B across both task groups.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.