acceptodds
Under review as a conference paper at ICLR 2027

Hard Preference Calibration for On-Policy Distillation

Abstract

On-policy distillation (OPD) has emerged as an effective paradigm for post-training large language models, yet teacher preferences can conflict with verifier outcomes. In practice, a teacher may prefer an incorrect trajectory over a correct one, making token-level guidance unreliable on critical examples. We thus propose Hard Preference Calibration (HPC), which augments OPD with an adaptive contrastive objective that calibrates verifier-derived hard preferences using the teacher's pairwise reliability. For each correct-incorrect trajectory pair, HPC applies stronger correction as the teacher's preference becomes less trustworthy, while preserving OPD's dense supervision. Our theoretical analysis further shows that this calibration yields a strictly tighter generation-error bound than OPD and improves Pass@k. Across multiple mathematical reasoning and code generation benchmarks, it outperforms OPD and related variants across student scales and teacher regimes, with absolute Pass@16 gains of up to 5.3% on HMMT25 and gains particularly evident when teacher preferences are less well aligned with verifier outcomes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.