Hard Preference Calibration for On-Policy Distillation
Abstract
On-policy distillation (OPD) has emerged as an effective paradigm for post-training large language models, yet teacher preferences can conflict with verifier outcomes. In practice, a teacher may prefer an incorrect trajectory over a correct one, making token-level guidance unreliable on critical examples. We thus propose Hard Preference Calibration (HPC), which augments OPD with an adaptive contrastive objective that calibrates verifier-derived hard preferences using the teacher's pairwise reliability. For each correct-incorrect trajectory pair, HPC applies stronger correction as the teacher's preference becomes less trustworthy, while preserving OPD's dense supervision. Our theoretical analysis further shows that this calibration yields a strictly tighter generation-error bound than OPD and improves Pass@k. Across multiple mathematical reasoning and code generation benchmarks, it outperforms OPD and related variants across student scales and teacher regimes, with absolute Pass@16 gains of up to 5.3% on HMMT25 and gains particularly evident when teacher preferences are less well aligned with verifier outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.