acceptodds
Under review as a conference paper at ICLR 2027

Distill the Teacher, Credit the Action: What On-Policy Distillation Transfers to a Visual Agent's Actions

Abstract

On-policy distillation (OPD) densely supervises student-generated trajectories, but for agents that act before answering, an action is distilled before execution reveals whether it was useful. For visual crop agents, we find that nearly all of OPD's gain comes from learning to read the acquired crop rather than where to look. Meanwhile, sampled crops shift toward the teacher's action distribution, which in our main setting is less aligned with the evidence. OPD uses each executed crop to train the reader but never credits the action for what that crop reveals—an *action-credit gap*. Execution-Conditioned Entropy Drop (Exec-ED) closes this gap without labels or changes to OPD: it credits each sampled action by how much its executed crop reduces the teacher's answer uncertainty, centered among sibling crops sampled for the same state. Across two teacher sizes and two benchmarks, Exec-ED improves crop acquisition while maintaining comparable end-to-end accuracy. In the main 9B setting, it raises ZoomBench crop precision from 55.0% under OPD to 57.4–58.8% (Base: 59.2%), and its crops are more useful to a held-out reader. This weak credit works in tandem with OPD's action KL: the KL anchors the actor near the teacher, while Exec-ED tilts its sampled actions toward more useful crops. More broadly, a teacher can show a student what to do, but only execution shows what worked; distillation for agents that act should learn from both.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.