acceptodds
Under review as a conference paper at ICLR 2027

When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation

Abstract

Reinforcement learning with verifiable rewards provides a ***sparse*** post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a scalar-weighted teacher KL term to the policy-gradient objective, providing ***dense*** token-level guidance that may be unreliable at some positions. Despite the benefits of combining these signals, their interaction during optimization can destabilize joint training. To understand how this instability develops, we study the learning dynamics of hybrid reward–distillation training through a neural tangent kernel (NTK) analysis. We introduce the ***cross-signal NTK*** , a token-level statistic that measures the alignment between reward and distillation gradients at position . Through this analysis, we identify two failure modes: 1 ***Magnitude drowning***, where the reward gradient exceeds the distillation gradient by orders of magnitude, so that even weak directional conflict can cause the distillation loss to ***rise*** despite its explicit inclusion in the training objective; and 2 ***Localized directional conflict***, where the sequence-level advantage and the teacher's position-specific distribution induce opposing updates at the same token (). The severity of these effects depends on the optimization regime: the gradient-norm ratio varies by roughly an order of magnitude across tasks, and our experiments reveal an empirical threshold beyond which naive mixing can lead to persistent training collapse. Motivated by these findings, we introduce the ***M3 family***, which combines magnitude normalization with three strategies for coordinating dense teacher supervision and sparse reward updates: a hard NTK-based mask that retains compatible teacher signals (M3-Select), a continuous relaxation of this mask (M3-Soft), and a fast–slow extragradient step that temporally separates teacher shaping from reward correction (M3-EG). Experiments across four model backbones and four benchmarks show that M3 maintains stable training dynamics and achieves superior performance in high- regimes where scalar-mixing baselines collapse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.