acceptodds
Under review as a conference paper at ICLR 2027

Token-Adaptive Concrete Score Distillation (TA-CSD)

Abstract

Knowledge distillation for autoregressive language models commonly applies a fixed temperature transform to logits before computing a divergence. We begin with an analytical observation: for any objective that depends on logits solely through pairwise differences, including softmax divergences over a fixed candidate window and concrete-score objectives, the centering component of logit standardization is mathematically inert. Consequently, standardization operates entirely as a per-token temperature , collapsing the question of whether to standardize into which estimator of that per-token scale to use. The conventional plug-in choice, the student's head standard deviation , is blind to teacher confidence, and because concrete-score gradients scale inversely with temperature, plug-in standardization silently acts as an effective learning rate modifier. We empirically quantify this effect in language model distillation, offering a candidate explanation for conflicting reports on logit standardization in the literature. Token-Adaptive Concrete Score Distillation (TA-CSD) predicts the per-token temperature via a lightweight multi-layer perceptron (MLP) gate conditioned on detached supervisory and predictive uncertainty features. In practice, the gate is trained jointly with the student under score distillation with detached features, while bilevel soft-ECE optimization provides a theoretical and diagnostic foundation. Under its learned scale, TA-CSD remains effective-learning-rate neutral (measured ratio , with mean ), so its differences from unscaled distillation cannot be attributed to step size. Evaluating across multiple model families and seeds with paired example-level bootstrap tests reveals a dissociation between calibration error and task capability. Plug-in standardization lowers top-1 calibration error only nominally ( vs. for CSD, not significant) while hedging (flattening the predictive distribution), losing of decisive () token predictions, degrading mathematical reasoning on GSM8K by percentage points (), and significantly worsening proper score (). While on-policy rollouts (GKD) achieve strong calibration, they alter the training distribution and demand over more training compute. Among offline cached distillation methods, TA-CSD avoids this harm: our clamped formulation attains the best point-estimate Brier score () and protects GSM8K reasoning (, vs. for SWS), while a residual gate reaches smECE at the cost of reasoning ().

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.