acceptodds
Under review as a conference paper at ICLR 2027

CoRD: On-Policy Distillation via Correction Reconstruction

Abstract

On-policy distillation helps language models improve reasoning through teacher feedback on their own responses, but selecting fewer supervised tokens can distort the collective learning signal. We aim to preserve this signal under a limited token supervision budget by accounting for how individual corrections reinforce or cancel one another. We introduce Correction Reconstruction Distillation (CoRD), which preserves the contributions of repeated correction directions to the average update. CoRD selects complementary token corrections in an output-head gradient representation and assigns signed weights to reconstruct the dense mean within their span, with theoretical bounds on approximation error in this representation. Across three student models and six mathematical reasoning benchmarks, including MATH-500 and AIME 2024/2025, CoRD improves macro mean@16 by 1.2–1.3 percentage points and pass@8 by 1.0–2.3 points over the strongest compared baselines under matched GPU-hour caps, using a 25% supervised-token budget. Coupling complementary selection with mean reconstruction allows CoRD to exploit redundancy while accounting for omitted tokens' aggregate contributions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.