acceptodds
Under review as a conference paper at ICLR 2027

Mixture RAG-KD: Controlling Retrieval-Induced Shifts for Robust Knowledge Distillation

Abstract

Knowledge distillation (KD) transfers the capabilities of large language models (LLMs) to smaller student models, but the limited parametric capacity of these students can constrain achievable performance. We argue that students should be distilled not only to imitate a teacher's parametric behavior, but also to effectively utilize external knowledge through retrieval-augmented generation (RAG). To this end, we formulate RAG-aware Knowledge Distillation, a framework that integrates retrieval into KD so that student models can effectively utilize information from retrieved contexts. We first identify a limitation of Vanilla RAG-KD, which directly distills the teacher's retrieval-conditioned distribution: under imperfect retrieval, its effectiveness is sensitive to retrieval quality. To address this issue, we propose Mixture RAG-KD, which combines the complementary strengths of the teacher's parametric and retrieval-conditioned distributions. We theoretically show why mixture supervision can be preferable to relying on either teacher alone by characterizing conditions under which an intermediate mixture achieves lower target risk than both endpoints. Building on this analysis, we develop global and instance-wise strategies for selecting the mixing coefficient by minimizing forward-KL or squared- target risk. Experiments across general QA and diverse noisy-retrieval settings show that Mixture RAG-KD consistently improves over Vanilla RAG-KD and remains more robust as training-time retrieval noise increases.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.