PAINT: Adaptive Privileged-Context Self-Distillation for Reasoning
Abstract
Effective reasoning supervision must act on the states a model visits while offering richer guidance than a final-answer reward. We frame on-policy self-distillation as an information-allocation problem: training-only context re-scores student-generated prefixes, and its influence is set by how much information is exposed and where the resulting token targets are calibrated. We propose PAINT (Privileged-context Adaptive INterpolated Training), which makes these two choices at different resolutions. For mathematical reasoning, rollout-reference overlap determines how much of a verified solution the privileged scorer sees; a teacher/student entropy ratio then selects a sparse set of prefixes for small energy-space interpolation. Conditional mutual information explains the value of context exposure, and geometric interpolation controls the local target. Across three competition-level math benchmarks and three Qwen3 scales, PAINT improves over full-solution on-policy self-distillation. On Qwen3-8B, it raises macro Avg@12 by 2.1 points over that baseline and 2.9 points over GRPO.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.