acceptodds
Under review as a conference paper at ICLR 2027

Learn from Whoever Is Right: Per-Sample Answer Verification for Multi-Teacher On-Policy Distillation

Abstract

Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable model remains challenging. By routing each sample to the teacher whose domain matches it, existing approaches let a domain label decide which teacher provides supervision. However, domain expertise holds only on average: the matched teacher is not always correct on a given sample, while a teacher from another domain sometimes is. The reliable teacher therefore has to be identified per sample, not per domain. In this paper, we introduce **M**ulti-**T**eacher **S**elf-**D**istillation **P**olicy **O**ptimization (MT-SDPO), an on-policy distillation method that unifies several frozen teachers into one student model. MT-SDPO consists of three components: (1) *self-anchors*, where a rollout is supervised by a correct rollout from its own group; (2) *answer-verified eligibility*, where a teacher may supervise a sample only if its own answer passes a verifier; and (3) *privileged distillation*, which merges the anchor and the feedback of all eligible teachers into one context that an exponential moving average self-teacher reads and the student does not, thereby keeping one policy at deployment. On three scientific domains with verifiable answers, MT-SDPO outperforms domain routing in worst-domain accuracy for all five students from three model families. On Qwen3-8B, it lifts worst-domain accuracy by 14.79 points and narrows its domain gap by 74.7%, a better balance than serving one matched teacher per domain. Verified reliability, not domain membership, should decide who teaches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.