EM-VLA: Closing the Loop Between Execution Monitoring and Action Generation for Long-Horizon Manipulation
Abstract
Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, yet reliable long-horizon execution remains challenging. Existing approaches introduce high-level reasoning or execution monitoring, but execution feedback is often used primarily to trigger intervention rather than continuously inform both reasoning and control. We propose EM-VLA, a hierarchical VLA that coordinates high-level reasoning and low-level control through a persistent state maintained by a lightweight execution monitor. The monitor integrates observed transitions, visual-language context, and internal action-model responses to update this state across interactions. The state guides subtask transitions and replanning while adapting action generation to the current execution stage. Internal responses from action generation then inform subsequent state updates, closing the loop between monitoring and control. This design enables the VLA to revise its plan and behavior as execution conditions change without requiring high-level reasoning at every action query. To support training, we derive temporally aligned subtask plans and supervision for reasoning and monitoring from successful demonstrations. Evaluations on LIBERO, CALVIN, RoboTwin 2.0, and real-world robotic tasks show improved long-horizon performance over the compared baselines and stronger recovery from execution-time perturbations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.