acceptodds
Under review as a conference paper at ICLR 2027

Progress-Aware On-Policy Distillation for Multimodal Agents

Abstract

Standard on-policy distillation (OPD) applies a uniform objective across all states. In long-horizon multimodal tasks, however, local alignment with a teacher often decouples from actual student progress. We observe that this uniform supervision frequently traps compact agents in low-progress continuation—they repeatedly invoke tools or search without meaningfully advancing toward the final answer. We present -OPD, a simple yet effective progress-aware distillation method. Our core idea is to condition teacher supervision on the student's realized progress. We quantify this progress via the step-wise change () in answer readinessin \emph–how ready the student is to answer correctly at the current step. -OPD dynamically adapts the learning objective: it applies reverse KL for interactions that actively drive progress, while switching to forward KL for low-progress steps to explore teacher-supported alternatives. Furthermore, for verified-success trajectories, it naturally bypasses redundant interactions where readiness is already saturated. Importantly, -OPD introduces zero inference overhead. It operates entirely on student-generated trajectories without requiring additional rollouts. Evaluated on the OmniGAIA benchmark, -OPD improves the accuracy of a Qwen3.5-4B student from 40.8% under standard OPD to 45.4%, while reducing the budget-exhaustion rate by 35.9%. Ablation studies show that selecting where to apply forward KL matters beyond its overall frequency, while readiness-based masking provides further gains. Code is provided in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.