EVIDENCE-ADAPTIVE DUAL-SPACE DISTILLATION FOR LANGUAGE MODELS
Abstract
Cross-tokenizer knowledge distillation requires both alignment between hetero- geneous output spaces and an appropriate allocation of supervision across them. We introduce Evidence-Adaptive Knowledge Distillation (EAKD), which uses detached student and teacher predictive cross-entropies to adapt the directional co- efficients of an existing dual-space objective. A learnable convex fusion preserves their common loss level, while a fixed absolute-gap term provides a conservative upper bound against local redistribution of the fusion coefficients. We formulate supervision allocation as a loss-based generalized-Bayesian decision problem: under specified endpoint actions, the unique control distribution and its posterior- mean action recover the implemented complementary directional weights exactly. The controller is recomputed on every mini-batch, preserves a fixed coefficient budget, and requires no additional language model. Our primary evaluation uses two LoRA-based TinyLLaMA-1.1B distillation settings with a fixed distillation temperature and dynamic weight allocation, improving average ROUGE-L over DSKD by 2.04 points for Mistral-7B → TinyLLaMA-1.1B and 2.08 points for LLaMA2-7B → TinyLLaMA-1.1B. The code will be released soon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.