acceptodds
Under review as a conference paper at ICLR 2027

CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation

Abstract

On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but typically treats response positions uniformly. Selective OPD addresses this limitation by allocating supervision non-uniformly across response tokens according to their estimated training value. Existing criteria, however, focus primarily on optimization need, such as uncertainty or teacher–student disagreement, while task relevance—whether supervision is tied to the semantic content of the current input—remains less directly characterized. We introduce Counterfactual Relevance for On-Policy Distillation (CROP), which operationalizes task relevance through a paraphrase-calibrated counterfactual sensitivity margin. For each source prompt, CROP constructs a validated original–paraphrase–counterfactual triplet, holds the student rollout fixed, and ranks response positions by how strongly the teacher reacts to a task-relevant condition change relative to a meaning-preserving paraphrase. At a 20% per-response token budget, CROP achieves the best aggregate performance in two teacher–student settings, outperforming the strongest non-CROP baseline by 1.43 and 0.96 points, respectively. Across 5%–20% supervision budgets, CROP consistently outperforms the strongest matched hard selector, with the largest observed margin at the 5% budget. These results suggest that task relevance provides a complementary signal for selective on-policy distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.