CoRD: On-Policy Distillation via Correction Reconstruction
Abstract
On-policy distillation helps language models improve reasoning through teacher feedback on their own responses, but selecting fewer supervised tokens can distort the collective learning signal. We aim to preserve this signal under a limited token supervision budget by accounting for how individual corrections reinforce or cancel one another. We introduce Correction Reconstruction Distillation (CoRD), which preserves the contributions of repeated correction directions to the average update. CoRD selects complementary token corrections in an output-head gradient representation and assigns signed weights to reconstruct the dense mean within their span, with theoretical bounds on approximation error in this representation. Across three student models and six mathematical reasoning benchmarks, including MATH-500 and AIME 2024/2025, CoRD improves macro mean@16 by 1.2–1.3 percentage points and pass@8 by 1.0–2.3 points over the strongest compared baselines under matched GPU-hour caps, using a 25% supervised-token budget. Coupling complementary selection with mean reconstruction allows CoRD to exploit redundancy while accounting for omitted tokens' aggregate contributions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.