Evolution-Aware Distillation via Optimal Transport for Large Language Models
Abstract
Knowledge distillation (KD) has become a widely adopted approach for compressing large language models (LLMs) for efficient deployment and inference. Existing LLM distillation methods can be broadly categorized into black-box and white-box paradigms. Black-box distillation transfers teacher-generated outputs but cannot fully exploit internal knowledge, while white-box distillation leverages logits or intermediate representations for finer-grained knowledge transfer. However, existing intermediate distillation methods usually rely on static layer-wise alignment and still suffer from two major limitations. First, predefined layer correspondence ignores that teacher and student predictions may evolve at different rates across model depth, causing layers at mismatched evolution stages to be aligned. Second, even with appropriate correspondence, indiscriminately transferring teacher evolution signals overlooks whether such knowledge actually complements the deficiencies of the student model. To address these issues, we propose an **ev**olution-**a**ware **k**nowledge **d**istillation framework (**EvaKD**) via optimal transport for LLMs, which dynamically aligns prediction evolution between teacher and student models and selectively transfers beneficial knowledge. Specifically, first, to overcome rigid layer alignment, we introduce prediction-evolution optimal transport, which projects intermediate representations into the prediction space, characterizes depthwise evolution through prediction changes between adjacent states, and establishes a soft many-to-many correspondence between teacher and student evolution via optimal transport. Second, to avoid indiscriminate knowledge transfer, we further propose deficiency-aware truncated distillation, which evaluates transported teacher evolution according to the remaining prediction deficiency of the student and retains only transport signals that contribute to closing this deficiency. Extensive experiments across multiple benchmarks and LLM families show that our method consistently outperforms strong LLM distillation baselines, improving Rouge-L scores by up to 7.01% and accuracy scores by 7.4%. Our code is at https://anonymous.4open.science/r/EvaKD-CFCF/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.