TBRD: Task-Aligned LLM Distillation via Brier Relations over Candidate Tokens
Abstract
Knowledge distillation transfers knowledge from a teacher model to a student by matching their output distributions. However, distribution matching supervises candidate preferences implicitly. Its updates can also conflict with the learning signal from the ground-truth token. We propose TBRD to explicitly model candidate relations and adapt the strength of teacher supervision. TBRD includes the teacher's highest-probability token in the student candidate set, models pairwise preferences, and adjusts the weight at each position using the agreement between distillation and task gradients. TBRD improves the average score across nine evaluation tasks by 2.09% over TAD, the strongest baseline in this comparison. These results support jointly modeling candidate relations and adapting supervision to the student's task signal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.