Junctions and Basins: Trajectory-Sensitive On-Policy Distillation for Language Models
Abstract
On-policy distillation (OPD) is an effective paradigm for transferring knowledge between language models, where a student learns from dense token-level supervision along its own trajectories. However, existing OPD methods typically apply uniform supervision across token positions, overlooking the substantial variation in their downstream consequences. We find that student trajectories exhibit two characteristic regimes: Junctions, where local variations strongly influence subsequent generation, and Basins, where plausible alternatives induce only limited changes to the future trajectory. To identify these regimes, we introduce a lightweight gradient-based future-sensitivity estimator that measures the directional sensitivity of the sampled future trajectory at each token position, together with a two-component mixture model that adaptively captures low- and high-sensitivity regions. Building on this structure, we propose Junction–Basin On-Policy Distillation (JB-OPD), which applies mode-seeking distillation at Junction-like positions while using mode-covering supervision with relational preservation at Basin-like positions. Across six mathematical reasoning benchmarks, JB-OPD consistently outperforms the state-of-the-art OPD baseline across Qwen3-0.6B, 1.7B, and 4B students, with average Pass@8 gains of 3.65%, 4.46%, and 4.39%, respectively. Further analysis reveals clear heterogeneity in downstream influence and shows that adapting distillation according to the inferred Junction–Basin roles consistently improves reasoning performance, supporting future trajectory sensitivity as an effective signal for allocating on-policy supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.