CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation
Abstract
On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but typically treats response positions uniformly. Selective OPD addresses this limitation by allocating supervision non-uniformly across response tokens according to their estimated training value. Existing criteria, however, focus primarily on optimization need, such as uncertainty or teacher–student disagreement, while task relevance—whether supervision is tied to the semantic content of the current input—remains less directly characterized. We introduce Counterfactual Relevance for On-Policy Distillation (CROP), which operationalizes task relevance through a paraphrase-calibrated counterfactual sensitivity margin. For each source prompt, CROP constructs a validated original–paraphrase–counterfactual triplet, holds the student rollout fixed, and ranks response positions by how strongly the teacher reacts to a task-relevant condition change relative to a meaning-preserving paraphrase. At a 20% per-response token budget, CROP achieves the best aggregate performance in two teacher–student settings, outperforming the strongest non-CROP baseline by 1.43 and 0.96 points, respectively. Across 5%–20% supervision budgets, CROP consistently outperforms the strongest matched hard selector, with the largest observed margin at the 5% budget. These results suggest that task relevance provides a complementary signal for selective on-policy distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.