Locality Guided Self-Distillation
Abstract
On-policy self-distillation (OPSD) trains a language model on its own responses using supervision from the same model given privileged, training-only context. However, this privileged context can improve reasoning while also altering response style, so matching the teacher’s full predictive distribution may reinforce changes unrelated to the desired reasoning improvement. We introduce the **Target-Locality Hypothesis**, with two testable predictions. **P1 (Correction Retention):** distillation targets that remain local to the student’s predictive distribution can retain most of the useful corrections in the privileged proposal. **P2 (Preference Advantage):** learning from such local targets can yield responses that align better with human preferences than either the unadapted student or full distillation. Motivated by this hypothesis, we propose **Locality-Guided Self-Distillation (LGSD)**, which constructs a distillation target that is as close as possible to the privileged proposal while remaining within a response-level KL budget around the current student, and then trains the student on this local target. Empirical studies, including human-preference evaluation on Arena, provide evidence for both P1 and P2. Across mathematical and logical reasoning benchmarks, LGSD improves over matched OPSD and other baselines, while additional analyses show improved training stability and reduced style drift.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.