acceptodds
Under review as a conference paper at ICLR 2027

On-Policy Distillation of Language Models by Repairing the Future

Abstract

On-policy distillation (OPD) supervises student-generated trajectories, but local token corrections do not guarantee better final outcomes because they can redirect the entire autoregressive continuation. We propose REPAIR-OPD, which identifies candidate repair points from teacher confidence and teacher–student disagreement, applies only a local teacher correction, and lets the student regenerate the remaining trajectory. Only Judge-verified wrong-to-correct repairs are used for learning. Across five mathematical-reasoning and two out-of-domain benchmarks, REPAIR-OPD achieves an overall mean of 54.55%, outperforming the strongest OPD baseline by 4.72%, and remains effective across smaller Qwen and cross-family Llama settings. In a matched one-round efficiency test on 1,000 source questions, REPAIR-OPD requires 1.36× OPD's end-to-end training time, while its seven-benchmark mean response length reaches 1.69× that of OPD, indicating a controlled cost–performance trade-off. Our results suggest that effective OPD should optimize not only local corrections, but the future trajectories they induce. Code is available at [the anonymous repository](https://anonymous.4open.science/r/repair-opd-code-2027-021D/README.md).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.