Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation
Abstract
Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, however, select one teacher per example and keep it fixed throughout the response. This design both depends on domain labels that mixed training corpora often lack and cannot adapt teacher selection when the expertise required changes within a trajectory. We observe that each specialist deviates more from a shared reference on in-domain prompts than on out-of-domain prompts, on average. Based on this observation, we propose TrustMOPD, which replaces example-level teacher selection with label-free, token-level supervision allocation. At each student-generated prefix, TrustMOPD measures this displacement in next-token preferences, calibrates its magnitude across teachers, and uses the resulting scores as proxies for local reliability to weight teacher-specific distillation losses. Evaluated across mathematics, code, and instruction following, TrustMOPD closes 91.5% and 98.0% of the overall-score gap between the initial student and oracle-routed teachers when trained on SingleCap and MultiCap, respectively, compared with 54.4% and 54.5% for the strongest label-free baseline in each setting. On SingleCap, it approaches label-based MOPD without using domain labels. Our code is available at https://anonymous.4open.science/r/TrustMOPD-F624.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.