acceptodds
Under review as a conference paper at ICLR 2027

When and Where Should Teachers Stop Teaching? Local Discriminability Collapse in Strong-to-Weak On-Policy Distillation

Abstract

In strong-to-weak on-policy distillation (OPD), where a high-capacity teacher supervises a substantially smaller student on student-generated rollouts with dense token-level feedback, conventional practice assumes that supervising every generated token is uniformly beneficial whenever teacher guidance is accessible. In this paper, we uncover a prevalent and trajectory-specific failure mode that overturns this premise: local discriminability collapse. Specifically, while the global teacher-student advantage (the mean log-probability gap on sampled tokens) persists late into a response, the teacher's local discriminative contrast across the student's reachable alternative actions sharply decays, forcing standard full-trajectory OPD to optimize over low-signal, high-variance suffix noise. To resolve this mismatch, we introduce Local-Discriminability-Aware On-Policy Distillation (LDA-OPD), a framework that selectively restricts dense teacher supervision to trajectory regions where guidance remains locally discriminative. Concretely, LDA-OPD applies three steps to each trajectory: it evaluates the teacher's log-probability margin between the top-two competitors within the student's high-probability candidate set, aggregates these margins over sentence-level semantic segments, and identifies the downward transition using a lightweight change-point test based on the profiled Bayesian Information Criterion (BIC) that is sweep-free over cutoff lengths. Past this boundary, we release dense supervision by zeroing the token-level loss on the suffix, and preserve per-sample loss mass by rescaling retained-prefix advantages so that the total loss magnitude remains unchanged. This adaptively truncates uninformative suffixes in a single training pass without requiring hyperparameter sweeps over cutoff lengths, while leaving student inference-time generation unconstrained. Extensive evaluations across multiple student scales (1.7B, 4B, 8B), model families (Qwen3, Gemma3), and diverse task types (multi-step mathematical reasoning, programmatic code generation, and out-of-domain scientific reasoning and instruction following) show that LDA-OPD consistently outperforms standard full-trajectory OPD and matches or surpasses multi-run fixed-prefix oracles in a single training run. Furthermore, matched-framework signal ablations and multi-seed evaluations provide strong empirical and counterfactual evidence that these gains are driven by local action contrast rather than generic entropy or position heuristics, establishing that effective strong-to-weak distillation requires evaluating not merely the presence of teacher guidance, but its local discriminative sharpness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.