acceptodds
Under review as a conference paper at ICLR 2027

Not All Tokens Should Update the Same Parameters: Prefix-Conditioned Gradient Routing for On-Policy Distillation

Abstract

On-policy distillation (OPD) trains a student on its own rollouts with token-level feedback from a stronger teacher. Because this feedback is issued on student-generated prefixes, it can also push the student away from behavior it already performs correctly, causing regressions on problems it previously solved. Existing remedies reweight tokens, but a scalar weight shrinks every parameter component of a token gradient together, so suppressing the part of an update that conflicts with previously successful behavior also discards the part that agrees with it. We propose Prefix-Conditioned Gradient Routing (PCGR), which makes this choice per parameter block. Using the gradient of the student's own verified past successes as a reference, PCGR lets the accumulated teacher–student discrepancy of the prefix decide how much reference conflict an update may keep, and blockwise reference geometry decide where the rest is removed. The resulting convex problem has a closed-form gate rule that, at equal retained reference conflict, never distorts the OPD gradient more than the corresponding group-level scalar gate. In standard OPD, 71.2% of token gradients mix reference-compatible and reference-conflicting blocks. Across four verifiable task families with Qwen3-1.7B and Qwen3-4B students, PCGR improves the macro score over OPD by 2.5 and 2.0 points and nearly halves the regression rate on previously solved problems (13.9% to 7.7%) without lowering the rate of new solutions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.