acceptodds
Under review as a conference paper at ICLR 2027

When Selective On-Policy Dstillation Rules Conflict: From Supervision Collapse to Minimal Repair

Abstract

On-policy distillation (OPD) uses teacher feedback to train a student on responses generated by its current policy. Existing methods select which tokens contribute to the training loss, using criteria such as uncertainty or teacher–student disagreement. However, these methods can select different positions, leaving too little shared feedback when every retained token must satisfy two rules. We discover that strict combination can eliminate all teacher feedback during token selection, yielding a zero joint gradient even when each rule alone produces a non-zero gradient. In this paper, we introduce SupportGuard to restore feedback through a weighted loss on tokens selected by only one rule. Separate training runs vary the mean loss weight over all response tokens. The lowest tested level meeting the parameter update criterion defines the repair target. SupportGuard compares the current mean loss weight with this target. Combinations that meet the target remain unchanged. Otherwise, it retains the full loss contributions of jointly selected tokens and applies a common weight to losses on tokens selected by only one rule. This weight equals the gap to the target divided by the increase in mean weight when these tokens receive full weight. We evaluate the mechanism and repair across Qwen-MATH, Granite-MATH, and Qwen-Code. Across the tested unsafe settings, SupportGuard restores parameter updates using, on average, only 1.02% of the additional training weight available from strict combination to the union of tokens selected by either rule. It leaves the safe control unchanged and improves accuracy on the fixed MATH evaluation by 2.11 percentage points over strict combination.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.