Training-free Multi-teacher Distillation as a Continual Learning Problem
Abstract
Multi-teacher distillation (MTD) has achieved notable success in integrating diverse capabilities into a single model; however, existing approaches suffer from immense data and computational costs. While model merging offers a training-free alternative, existing techniques largely rely on empirical heuristics and lack a rigorous theoretical foundation. In this paper, we novelly formulate MTD as a continual learning problem. Specifically, we conceptualize independent teachers fine-tuned from a shared base model as having been trained sequentially, where acquiring a "new" task induces catastrophic forgetting on "previous" ones. Consequently, we propose **Distill**-by-**Pro**jection (Distill-Pro), a post-hoc parameter rectification technique that projects and corrects model updates to eliminate components that interfere with previous task performance. Crucially, Distill-Pro remains entirely data-free and training-free, while being grounded in a theoretical foundation of first-order performance invariance. Extensive evaluations on two large language model multi-teacher distillation benchmarks—encompassing two model architectures and three diverse tasks—demonstrate that our method consistently outperforms or matches state-of-the-art model merging and on-policy distillation baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.