acceptodds
Under review as a conference paper at ICLR 2027

HEIR: When the Student Inherits the Teacher's Seat Pareto-Competitive Reference Selection for On-Policy Self-Distillation of Reasoning LLMs

Abstract

On-policy self-distillation improves reasoning models by training on student rollouts while a teacher receives a privileged reference solution. That reference is usually fixed, even after the student begins producing better teaching examples. We introduce HEIR (Hint-Earning by Iterative Rollouts), a simple reference-handoff rule: sample a group of student rollouts, verify them, and replace the reference only when the shortest correct rollout clears a relative length margin. The admitted rollout conditions the teacher and supplies the student trajectory; the distillation objective itself is unchanged. Across Qwen3-1.7B/4B/8B and three competition-math benchmarks, HEIR improves a tuned fixed-reference baseline on every 4B and 8B Avg@12 result, with gains up to 3.34 points while shortening responses. A matched 4B comparison separates group selection from reference handoff: winner selection contributes 1.82 macro-average points over standard distillation, and handoff adds another 0.87 points (95% CI: 0.31–1.38) at essentially identical cost. The handoff gain is negligible for compact references but grows to 1.45 points for padded ones; it remains positive after non-verbatim and answer-masked controls and under injected verifier false positives. These results identify the operative regime precisely: a capable student, a verifiable task, and a reference whose form has become a constraint rather than an asset.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.