Selective Feedback Attenuation for Multi-Teacher On-Policy Distillation
Abstract
Multi-teacher on-policy distillation consolidates specialized capabilities into a single student by leveraging teacher feedback on student-generated responses. However, strictly matching teacher preferences can penalize valid alternative solutions, creating a fundamental misalignment between task correctness and imitation signals. To address this, we propose Selective Attenuation for Multi-Teacher On-Policy Distillation (SA-MOPD). SA-MOPD uses verifier outcomes to attenuate negative feedback on successful student responses, while preserving positive feedback and standard supervision elsewhere. The rule modifies feedback coefficients within the existing distillation update. Our reverse-KL analysis identifies a mechanism by which disagreement among successful solutions can reduce student success despite a stronger teacher. A local analysis further characterizes when feedback selection improves task performance at a fixed attenuation budget. Our experiments demonstrate the value of selective supervision retention for capability integration: SA-MOPD raises recovery from Open-MOPD's 83.4% to 96.4% across mathematical reasoning, code generation, and instruction following. These gains extend to agentic tasks, where the single student surpasses the routed teacher ensemble in aggregate performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.