acceptodds
Under review as a conference paper at ICLR 2027

Beyond Teachers: Student-Ascendant Knowledge Distillation for LLM-based Machine Translation

Abstract

Knowledge distillation (KD) can effectively compress a large machine translation (MT) model into a smaller variant while preserving performance, supporting deployment in constrained environments, especially under the context of large language model (LLM)-based MT. However, the smaller student model is intrinsically bounded by the performance ceiling of the teacher model in recent studies of KD. Yet for humans, students routinely outstrip their teachers through progressive, deliberate learning. In this paper, we rethink the process of KD and propose SaKD, tudent-scentant nowledge istillation, to break the performance ceiling of teacher model in KD for LLM-based MT via mimicking the human learning behaviors. SaKD first distills knowledge from multiple teachers, and then introduces a human-like three-stage learning process, , and —to guide the student model toward surpassing its teachers. Experiments on extensive MT benchmarks show the effectiveness of SaKD, significantly outperforming existing KD methods. And importantly, the student model with SaKD transcends the performance bound of the teacher model, achieving a complete performance surpass.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.