CoKD: Generating Student-Reachable Targets through Teacher–Student Collaborative Decoding
Abstract
Knowledge distillation (KD) is a standard approach for compressing large language models, yet its effectiveness critically depends on the quality of the targets used to supervise the student. Teacher-generated targets can be informative but often lie outside what a small student can reliably learn, whereas student-only rollouts remain reachable but may encode suboptimal decisions. We propose Collaborative Knowledge Distillation (CoKD), which generates distillation targets through teacher–student collaborative decoding. At each decoding step, CoKD samples from a tempered distribution that combines teacher preference with student support, favoring tokens that are endorsed by both teacher and student. A task-aware sequence-level quality filter further retains only collaborative responses that provide useful training targets. Across arithmetic reasoning, instruction following, and dialogue summarization benchmarks with Qwen2.5 and Gemma teacher–student pairs, CoKD consistently improves over competitive KD baselines and remains stable across varying teacher scales. These results indicate that effective distillation requires targets that are not only teacher-preferred but also student-reachable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.