acceptodds
Under review as a conference paper at ICLR 2027

SELD: Selective On-Policy Distillation for Mathematical Reasoning via Semantic Error Localization

Abstract

On-policy distillation (OPD) trains students on their own solutions, but how should teacher supervision be allocated within failed reasoning? We investigate concentrating supervision near the first error that remains uncorrected and leads to the wrong answer. In a controlled placement study, a fixed supervision window centered on the diagnosed error outperforms the tested preceding, following, and random windows. Yet teacher–student disagreement need not identify that error. We propose Selective On-Policy Distillation (SELD), which uses a frozen teacher and a reference solution to diagnose the first error and organize supervision around it. SELD emphasizes the diagnosed step and its predecessor, retains sparse earlier-prefix supervision, and allocates nonuniform weights within these regions while excluding the later continuation. On 159 expert-annotated incorrect solutions, semantic diagnosis raises first-error coverage from 44.7% with maximum-disagreement anchors to 67.9%, evaluating both anchors with SELD's predecessor-plus-anchor transition. Across four mathematical reasoning benchmarks and three training seeds, SELD improves Qwen3-1.7B macro-average response accuracy by 2.54 percentage points over full-token OPD, with gains also observed for a larger student and a smaller teacher. Repeated ablations support the error-step, prefix, and weighting choices.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.