Before the KL: Auditing Interfaces for Cross-Tokenizer Distillation
Abstract
Cross-tokenizer knowledge distillation (CTKD) applies its loss not directly to native model distributions, but to virtual distributions constructed by an alignment and scoring interface. We show that matching token coordinates is insufficient to reproduce this objective: representations that preserve every native softmax probability can still change candidate scores, virtual distributions, and parameter gradients. For fixed-support linear scoring rules, we derive a necessary and sufficient invariance condition: every candidate must assign the same total weight to each native position. A pinned public SimCT implementation violates this condition. Its published formula aggregates log probabilities, whereas its runtime aggregates raw logits with different position weights for spans and overlaps. In a frozen Falcon3-to-Pythia audit on 224 texts, converting logits to differentiable log probabilities changes native probabilities by at most , yet yields a median complete-gradient cosine of . A matched short-horizon diagnostic further produces distinct parameter paths, while raw-logit and formula-score evaluators prefer different endpoints; native negative log-likelihood favors the raw-trained endpoint. These results show that formula fidelity, evaluator preference, and model quality are distinct quantities. CTKD implementations should therefore report the score construction and its gradient path, rather than alignment alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.