Learning from What the Teacher Trusts: Confidence-Aligned On-Policy Distillation
Abstract
On-policy distillation is a promising paradigm for aligning Large Language Models. However, standard objectives compel the student to indiscriminately mimic the teacher, forcing the model to learn from low-agreement or noisy trajectories generated during exploration. While recent approaches attempt to mitigate this via external reward models or step-wise teacher intervention, they introduce significant computational and communication overhead. In this work, we propose Teacher-Confidence Aligned On-Policy Distillation (TCAD), a scalable, rewardfree paradigm that leverages the teacher’s intrinsic sequence-level confidence as a relative ranking signal among rollouts for the same prompt. By dynamically reweighting trajectories based on model-relative teacher certainty, TCAD softly down-weights low-agreement rollouts without requiring auxiliary reward models. To address the approximation error caused by vocabulary truncation in large-scale distillation, we introduce support-symmetric candidate construction and probability-proportional Monte Carlo optimization over a compact event space. These mechanisms reduce the distribution mismatch between teacher and student supports while optimizing a compact support-plus-tail divergence with a controllable Monte Carlo sample-count trade-off. Extensive experiments across diverse benchmarks demonstrate that our proposed method consistently improves over standard on-policy distillation baselines, achieving strong alignment efficiency without external supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.