Distilling Productive Disagreement: Collision-Aware On-Policy Distillation for Self-Correcting Language Agents
Abstract
Teacher disagreement is useful for on-policy distillation only when it exposes a decision-sensitive alternative that can be verified and acted on from the student's current state. We study this conditional learning signal and introduce Collision-Aware On-Policy Distillation (C-OPD). For each student trajectory, an axiom teacher identifies implicit assumptions, a counterfactual teacher minimally inverts high-risk assumptions, and task-native verification adjudicates the resulting branches. A Productive Collision Score (PCS) combines disagreement, counterfactual utility, evidence consistency, recoverability, and training-trajectory reward improvement, then attributes additional distillation weight to an intervention-linked reasoning span. Across three seeds on SWE-bench Verified, BFCL v3.0.1, MATH-500, and HotpotQA v1.1, C-OPD reaches 49.6±0.4% macro-average success, compared with 41.8±0.4% for standard on-policy distillation and 38.7±0.4% for response-level knowledge distillation under the same 800M teacher-token budget per seed. On SWE-bench Verified, resolved instances increase from 32.4±0.4% to 38.9±0.3%. On adversarial tool observations, unsupported continuation falls by 28.4%, while recovery after an incorrect intermediate decision rises from 46.1% to 61.3%. Weighting only disagreement, evidence consistency, or future reward underperforms full PCS by 5.2, 4.4, and 4.8 points, respectively, showing that future reward alone does not explain the gain. Retraining without collision-generated supervision absent from all pre-collision teacher responses lowers success from 49.6% to 43.9%, accounting for 72.6% of the observed improvement over standard OPD under this controlled removal-and-retraining intervention. The results show that verified, recoverable disagreement supplies supervision that raw disagreement and best-of-teacher selection do not capture.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.