acceptodds
Under review as a conference paper at ICLR 2027

PACE: Prefix-Anchored Calibration with Adaptive Segmentation for On-Policy Distillation

Abstract

On-policy distillation (OPD) offers a practical way to align smaller student models with stronger teachers on student-generated trajectories. However, token-level teacher-student discrepancies are often noisy and poorly calibrated, as they are susceptible to capability-irrelevant noise, such as lexical variation, formatting artifacts and teacher preferences rather than reflecting differences in reasoning ability. Consequently, standard OPD struggles to distinguish meaningful teacher-student divergence from stable response-level mismatch. To address this problem, we propose PACE (Prefix-Anchored Calibration with adaptive sEgmentation), a segment-level calibration objective for on-policy distillation. PACE first adaptively partitions each student generated response into locally coherent segments using change points in teacher's negative log-likelihood (NLL), together with structural boundaries. PACE computes a student-to-teacher NLL ratio from segment-averaged losses, yielding a more stable local measure of mismatch than token-level discrepancies. The mean ratio over the initial segments serves as a trajectory-specific anchor against which subsequent segment ratios are calibrated. This relative formulation helps reduce the influence of persistent teacher–student discrepancies unrelated to reasoning ability, allowing the learning objective to focus more directly on segment-specific capability gaps. It complements standard OPD without requiring additional reward models or human annotations, providing a lightweight segment-level calibration signal for LLM agents. Experiments across tool-use and coding benchmarks demonstrate consistent improvements over OPD baselines across different teacher-student model scales, supporting the effectiveness and generality of PACE.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.