Learning to Verify: Allocating Revision Effort in Embodied Agents
Abstract
Improving an embodied agent requires recognizing when stable execution fails to satisfy an instruction. A robot may remain upright while moving in the wrong direction, traveling too far, or performing actions out of order. We propose Learn to Verify, then Revise (LVR), which turns feedback into executable task checks and uses them to guide subsequent revisions. LVR selects clarification questions by their expected effect on verification decisions, then reuses the learned checks without further feedback. Our analysis characterizes when a shared check is reliable enough for additional candidate search not to reduce expected task success. On recorded revision chains for 47 simulated robot tasks, LVR achieves 87.2% success, compared with 76.6% when every task is revised, while using 58% fewer revisions. With separately generated revision chains on 39 tasks, updated checks achieve 89.7% success versus 74.4% before feedback, selecting 50 versus 40 revisions under offline stopping. Classification and transfer tests show no corresponding established gain in held-out accuracy. These results motivate evaluating learned verification through the revision decisions it enables, alongside its predictive accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.