EMBER: Credit-Guided Distillation with Online Student Exploration
Abstract
Correct teacher traces do not specify which reasoning steps a changing student should imitate. We introduce Evolving Marginal-Benefit Estimation for Reasoning (EMBER), a framework that adapts supervision from fixed, text-only teacher traces to the current student's capabilities. EMBER estimates a step's credit by comparing terminal returns at adjacent teacher prefixes under the same student continuation policy. Positive credit weights imitation, while the sampled continuations also support reinforcement learning. Preserving credit magnitude allows teacher supervision to weaken as its estimated benefit declines. Our analysis characterizes credit reliability under estimation error and policy drift, and shows that question-only mastery alone does not make teacher credit vanish: value drops along teacher prefixes must also disappear. Experiments on instruction following and mathematical reasoning examine this coupling of guidance and exploration. On AMC 2022–2023, Qwen3.5-4B trained with EMBER exceeds SFT by 6.7 percentage points in avg@32, with training dynamics consistent with a diminishing teacher contribution. These findings support student-relative task return as a criterion for adapting teacher supervision during online learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.