acceptodds
Under review as a conference paper at ICLR 2027

Selective Feedback Attenuation for Multi-Teacher On-Policy Distillation

Abstract

Multi-teacher on-policy distillation consolidates specialized capabilities into a single student by leveraging teacher feedback on student-generated responses. However, strictly matching teacher preferences can penalize valid alternative solutions, creating a fundamental misalignment between task correctness and imitation signals. To address this, we propose Selective Attenuation for Multi-Teacher On-Policy Distillation (SA-MOPD). SA-MOPD uses verifier outcomes to attenuate negative feedback on successful student responses, while preserving positive feedback and standard supervision elsewhere. The rule modifies feedback coefficients within the existing distillation update. Our reverse-KL analysis identifies a mechanism by which disagreement among successful solutions can reduce student success despite a stronger teacher. A local analysis further characterizes when feedback selection improves task performance at a fixed attenuation budget. Our experiments demonstrate the value of selective supervision retention for capability integration: SA-MOPD raises recovery from Open-MOPD's 83.4% to 96.4% across mathematical reasoning, code generation, and instruction following. These gains extend to agentic tasks, where the single student surpasses the routed teacher ensemble in aggregate performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.