SAGE: Online Continual Knowledge Distillation for Small Language Models
Abstract
Small language models are increasingly deployed in resource-constrained settings because of their computational efficiency, yet their limited capacity often leads to performance degradation when they encounter previously unseen data. Although knowledge distillation can transfer capabilities from larger teacher models, conventional approaches are typically confined to offline training and therefore cannot provide continuous supervision after deployment. Here, we introduce SAGE, an online continual distillation framework that enables a deployed small language model to selectively acquire knowledge from a teacher model in streaming environments. SAGE comprises three coordinated modules. The Trigger identifies inputs beyond the student’s competence boundary by combining teacher–student discrepancies without requiring labels. The Buffer uses streaming clustering and stability filtering to organize triggered samples into semantically coherent knowledge regions. Finally, the Debugger performs a two-stage search over low-rank parameter updates to derive effective adaptation configurations for targeted knowledge injection. Experiments show that SAGE substantially improves both sample efficiency and adaptation accuracy, providing a practical framework for continuously enhancing small language models after deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.