acceptodds
Under review as a conference paper at ICLR 2027

LBSD: Local-Branch Selective Self-Distillation

Abstract

On-policy self-distillation (OPSD) mitigates the coarse-grained supervision issue of reinforcement learning with verifiable rewards (RLVR) with long reasoning trajectories by introducing dense token-level targets. However, standard OPSD methods largely treat all tokens alike, overlooking that only a subset of positions provides strong signals for exploration or correction. We introduce LBSD, a Local-Branch Selective self-Distillation framework that concentrates supervision on informative parts of the reasoning process. LBSD extends single-trajectory OPSD by selectively expanding critical intermediate states into local on-policy branches. It uses student entropy and teacher-student discrepancy to identify informative branches and token positions for learning. This design provides broader and more targeted feedback. Moreover, we propose abstract solution strategies as privileged teacher information instead of full reference reasoning trajectories, reducing the risk of answer leakage. Experiments across four competition mathematics benchmarks and multiple Qwen3 model scales demonstrate consistent improvements over representative RLVR and self-distillation baselines. Ablation studies further support the effectiveness of local branching and distribution-informed selection. These findings suggest that effective self-distillation depends not only on providing dense targets, but also on identifying where and along which local continuations those targets should be applied.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.