acceptodds
Under review as a conference paper at ICLR 2027

Spotter: Let the Embodied Model Lead, and the LLM Reflect for It

Abstract

Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt; this is the mechanism behind the gains language models obtain from thinking, where a model reflects on an error and revises. We test whether they can repair a known error, which requires producing a correction and judging whether it is right. When an external interrupt stops them at a failure and lets them retry, their own randomness seldom repairs the error, a language description of the error changes little, and best-of-N selection cannot pick the successful candidate in post-failure states. We attribute this to training on successful demonstrations, which contain no failure or recovery, and to inputs too narrow to show what went wrong, and conclude that reflection must be supplied by a vision-language model (VLM), which can take in far more information, such as the history of the episode and text, and is more general. Prior VLM-led work has the VLM plan every step and invoke the embodied model as a tool, placing the VLM on the critical path. We propose Spotter, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control. We run Spotter with Qwen and with GPT as the VLM, and both improve the embodied models; with GPT, Spotter improves Cosmos Policy and by 5.6 and 7.5 percentage points on RoboCasa, and raises from 47.2% to 57.0% on the Hard setting of RoboTwin 2.0 and from 53% to 83% on a real robot. Because the VLM steps in only when an error is confirmed, a successful episode with Qwen takes only 13 to 16 s longer than with the embodied model alone and about 70% less time than with a VLM-led baseline using the same model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.