acceptodds
Under review as a conference paper at ICLR 2027

ACT:Selective Temporal Feedback for On-Policy Distillation

Abstract

On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories, yet dense supervision can still be poorly allocated. Two positions with similar local teacher–student disagreement may have different trajectory contexts: the disagreement may quickly dissipate at one position but persist into the future at another. Exploiting this temporal context requires distinguishing where to intervene from what feedback to apply: position selection should preserve disagreement before sign cancellation, whereas feedback construction should retain its direction and net effect. Based on this distinction, we introduce Advantage Credit Transfer (ACT), which decouples position selection from feedback construction. Over a short future window, ACT uses unsigned future mismatch to select a sparse set of positions and mixes signed future feedback into the shaped directional coefficients at those positions, leaving all other shaped directional coefficients unchanged. ACT reuses the teacher and student log probabilities already computed by OPD, requiring no additional teacher queries, trajectory sampling, or value network. Controlled experiments show that future-mismatch selection outperforms alternative selectors at equal coverage, signed future feedback outperforms unsigned replacements under a fixed gate, and selective correction outperforms uniform allocation under matched total correction magnitude. With a 1.5B student, ACT achieves the highest macro average among the evaluated methods across five benchmarks spanning mathematical, coding, and scientific reasoning, under both math-domain and multi-domain distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.