Balancing Specialists with Safety: Lagrangian Anchored Multi-Teacher On-Policy Distillation for Large Language Models
Abstract
Multi-Teacher On-Policy Distillation (MOPD) merges diverse expert capabilities into a single student model. However, current methods rely on rigid routing or heuristics, making them highly vulnerable to multi-domain gradient conflicts. These unconstrained updates often cause a severe “see-saw” effect, where acquiring new skills catastrophically erodes previously established foundational capabilities. To address this, we propose Lagrangian Anchored Multi-Teacher Post-Training (**LAMP**), a framework that protects foundational knowledge by dynamically constraining policy updates. LAMP anchors the student to an initial reference model and employs a self-adjusting penalty mechanism: when aggressive specialist updates threaten to distort core capabilities, LAMP automatically scales up an anchoring penalty to pull the student back toward the reference distribution. Under explicit regularity and step-size assumptions, a theoretical analysis establishes subsequential primal stationarity and a Polyak-Lojasiewicz-based tracking rate. Experiments across model families with different architectures and scales show that LAMP improves aggregate multi-domain performance over strong MOPD baselines under different settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.