Breaking Sampled-Token Tunnel Vision Enables Robust Capability Transfer
Abstract
On-policy distillation (OPD) efficiently transfers capabilities acquired through costly post-training by providing dense teacher supervision along student-generated trajectories. Yet sampled-token OPD can become unreliable in cross-origin teacher–student settings. Through controlled comparisons between successful and failed training regimes, we identify sampled-token tunnel vision: analogous to human tunnel vision, the advantage estimator focuses on a single sampled token's teacher–student probability ratio while overlooking the update's effects on surrounding alternatives. This narrow view can obscure the suppression of teacher-preferred candidates or the amplification of less-preferred ones, undermining capability transfer. Simply aggregating reverse-KL contributions over nearby tokens can also lose directional information: the same aggregate may arise when the sampled token should be promoted in one case but suppressed in another, relative to its neighbors. We therefore propose **Neighborhood Contrastive On-Policy Distillation** (NC-OPD), which reconstructs the sampled-token advantage through pairwise comparisons within a lightweight neighborhood of tokens ranked above the sampled token by the teacher or below it by the student. This provides explicit relative-preference guidance while preserving the original sampled-token policy-gradient interface. Across cross-scale and cross-reasoning-mode transfer settings, **NC-OPD** consistently improves mathematical reasoning, code generation, and multi-domain performance. These results highlight neighborhood-relative advantage estimation as an effective research line to robust on-policy capability transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.