DCRD: Retrieval-Grounded Curriculum Distillation for Imbalanced Multi-Domain Question Answering
Abstract
Knowledge distillation has become an effective approach for transferring the capabilities of large language models into smaller and more efficient students. However, existing distillation methods typically assume a single teacher and balanced training distributions, which limits their applicability to realistic question-answering scenarios involving heterogeneous domains with severe data imbalance. We present DCRD, a multi-teacher distillation framework for imbalanced multi-domain question answering that treats each domain–teacher pair as a distinct learning course and dynamically schedules training effort based on the student's recent validation accuracy on each course. The framework combines retrieval-augmented knowledge grounding—enriching teacher supervision with external evidence from Wikipedia and PubMed—with a curriculum scheduler that adaptively adjusts course sampling and training intensity according to the student's performance. Experiments on MedQA-USMLE, OpenBookQA, and StrategyQA show that our approach consistently improves performance across Flan-T5 model scales, reaching 43.29% average accuracy on Flan-T5-Large and outperforming teacher ensemble and MoE-based distillation baselines without additional inference cost. These results demonstrate that curriculum-aware multi-teacher distillation provides an effective and scalable solution for knowledge-intensive multi-domain reasoning. Our code is available at https://github.com/linda-2-hh/DCRD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.