Every Fork Has a Price: The Precision–Coverage Trade-off in On-Policy Distillation of Reasoning Models
Abstract
How much of a teacher's reasoning diversity should a student keep? In on-policy distillation, a routing rule decides which tokens receive coverage transfer. The dominant answer, entropy-aware on-policy distillation (EOPD), applies forward KL wherever teacher entropy exceeds a threshold. We test whether smarter routing can do better. Five routing mechanisms of increasing sophistication are compared in a controlled design-space study under one strictly reproduced protocol (Qwen3-1.7B student, Qwen3-8B teacher, six math benchmarks, identical trajectory budgets). What survives is a principle: coverage transfer is budget-limited, not routing-limited. Mechanistically, gates conditioned on trajectory outcomes remove forward-KL mass precisely from low-success problems whose marginal Pass@k value is highest. Whole prompt groups lose all coverage signal at a rate matching a compounding miscoverage model ((1-p)^n). Budget-preserving reweighting cannot recover the coverage that closed prompt groups never receive. EOPD's crude threshold keeps the coverage lead because it spends the most: Pass@k follows the budget, not the gate. Three measurements make the principle concrete: a soft continuous entropy gate matches EOPD on macro Avg@8 within a ±3-point equivalence margin across seeds (29.99±0.46 vs. 30.15±0.54) at zero extra rollout cost. A cover-then-sharpen curriculum turns the budget into a dial: a half-schedule cover holds coverage level with the full-budget endpoint over three seeds (50.21±0.60 vs. 50.68±0.28 Pass@8). The forward-KL coefficient alone moves macro Avg@8 monotonically up to the default dose (28.54 → 30.63), confirming that outcomes track the budget, not the gate. Our code is available at the anonymous repository: https://anonymous.4open.science/r/TSA-OPD-for-ICLR-2027-50BD.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.