MetaOPSD: Adapting Privileged Teachers with Post-Update Student Feedback
Abstract
Access to reference solutions can make a language model a better predictor without making it a better teacher. On-policy self-distillation (OPSD) trains a problem-only student to match a reference-conditioned teacher on student-generated prefixes, but distributional agreement does not directly assess the learning gains induced by that supervision. We introduce MetaOPSD, which adapts the privileged teacher using verified student outcomes after a pilot distillation step. The pilot student solves separate feedback problems without privileged information, and a policy-gradient surrogate propagates outcome feedback through the pilot update to train the teacher. The adapted teacher then supervises an OPSD update from the original student state, retaining dense distillation as the student's training objective. MetaOPSD outperforms all evaluated baselines on each of three mathematical reasoning benchmarks at both model sizes after 200 student updates, with macro avg@12 gains of 2.16–2.96 percentage points over the OPSD baseline. On Qwen3-1.7B, checkpoint cross-evaluation further shows that the final teacher induces larger learning gains than the initial teacher at all evaluated student states.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.