CoR-OPD: Verified Local Repair and Handback for On-Policy Distillation of Long-Horizon Agents
Abstract
On-policy distillation (OPD) trains a student on its own trajectories, but an early mistake can derail a long-horizon agent. Teacher predictions on the resulting trajectory may be unreliable, and token-level matching along its unproductive suffix may not teach the student how to recover. We present Cor-OPD, which detects sustained teacher–student disagreement in decision tokens, returns to the last reliable action boundary, and requests a teacher bridge. With Qwen3-Coder-30B, Cor-OPD resolves 50.0% of SWE-bench Verified, compared with 45.8% for the student and 48.8% for standard OPD. With Qwen3.5-9B, it achieves 49.85% macro rubric completion on SWE-Atlas Codebase QnA, compared with 46.07% for the base student. It also transfers to multi-turn function calling and cross-language code editing. Behavioral analysis shows that bridge supervision improves evidence probing and schema repair, supporting verified local handback as a promising training strategy for long-horizon agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.