acceptodds
Under review as a conference paper at ICLR 2027

Learning to Verify: Allocating Revision Effort in Embodied Agents

Abstract

Improving an embodied agent requires recognizing when stable execution fails to satisfy an instruction. A robot may remain upright while moving in the wrong direction, traveling too far, or performing actions out of order. We propose Learn to Verify, then Revise (LVR), which turns feedback into executable task checks and uses them to guide subsequent revisions. LVR selects clarification questions by their expected effect on verification decisions, then reuses the learned checks without further feedback. Our analysis characterizes when a shared check is reliable enough for additional candidate search not to reduce expected task success. On recorded revision chains for 47 simulated robot tasks, LVR achieves 87.2% success, compared with 76.6% when every task is revised, while using 58% fewer revisions. With separately generated revision chains on 39 tasks, updated checks achieve 89.7% success versus 74.4% before feedback, selecting 50 versus 40 revisions under offline stopping. Classification and transfer tests show no corresponding established gain in held-out accuracy. These results motivate evaluating learned verification through the revision decisions it enables, alongside its predictive accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.