acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Imitate: Counterfactual Value Modeling for On-Policy Distillation

Abstract

On-policy distillation (OPD) addresses a central mismatch in reasoning distillation by querying the teacher on states actually visited by the student. Yet state alignment does not guarantee useful guidance: in multi-step reasoning, a locally valid teacher step may still lead to a worse final outcome under the student's continuation. Selective variants modulate imitation using uncertainty, disagreement, or teachability, but these proxies do not reveal whether substituting the teacher action improves the current student's eventual outcome. We introduce Counterfactual On-Policy Distillation (CFOPD) to estimate this policy-relative intervention value. CFOPD compares teacher and student actions through matched rollouts under the same frozen student policy and amortizes these comparisons with a shared counterfactual value model. The learned values induce two complementary contrasts: action-level contrasts guide supervision between teacher and student alternatives and modulate teacher distillation, while temporal contrasts track progress along the realized student trajectory. Because both depend on the student's continuation ability, the counterfactual supervision is refreshed as the student evolves. Across two Qwen3 teacher–student settings, CFOPD improves Avg@8 over GKD-OPD on all six reasoning benchmarks, with MATH-500 gains from 50.65 to 55.97 and 32.18 to 44.28, respectively. High-budget evaluation further finds the student action preferable at 30–50% of evaluated states despite the teacher's higher overall accuracy, supporting the need for policy-relative action evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.