acceptodds
Under review as a conference paper at ICLR 2027

Teaching Students Where to Look: Saliency-Guided Knowledge Distillation for LLMs

Abstract

Knowledge distillation (KD) is the standard approach for compressing large language models (LLMs) into compact deployable students by training the student to match the teacher's output distribution. While many improvements have been proposed, such as alternative divergence choices, on-policy generation, and difficulty-adaptive sample reweighting, existing methods deliver the distillation signal pointwise in output space, providing no constraint on how the student behaves in the neighborhood of training inputs. We empirically show that this pointwise objective leaves large teacher–student gaps in input neighborhoods even when output agreement is exact, and propose SaKD (Saliency-Guided Knowledge Distillation), which closes this gap by lifting pointwise output matching to a saliency-aware first-order matching objective. Concretely, SaKD combines a noise-injected Kullback–Leibler (KL) divergence term that, via a noise–Jacobian identity for autoregressive token-level KL, provides an implicit teacher–student Jacobian-matching signal along coupled perturbation directions, with a saliency-divergence reweighting that focuses the loss on samples with larger teacher–student input-saliency disagreement. Extensive experiments show that SaKD consistently outperforms seven strong KD baselines across Qwen3 and LLaMA student scales, with students reaching correct answers through teacher-like reasoning rather than surface shortcuts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.