Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-Aware Memory
Abstract
Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) and world-action models (WAMs) increasingly master the individual skills. However, the chain itself still fails: errors compound beyond the policy's ability to correct, and the execution of one subtask silently constrains the next. Coding agents offer a promising pathway: freeze the VLA and put an LLM agent in charge of it. The agent plans in language, moves in free space with analytic primitives, and invokes the VLA only for the contact-rich segments, writing all adaptation into language memory. Yet applied directly to long horizons, the agent-plus-VLA recipe breaks twice. 1) Its competence is acquired through whole-task exploration at test time, whose cost is exponential in the number of stages: where a single stage requires episodes on average, a -stage task requires on the order of , and a failed episode does not identify the stage responsible. 2) It has no representation of the transitions: the VLA primitive carries an exit condition but no entry condition, and a subtask can succeed in a form its successor cannot use. We present BATON, which addresses the two failures in turn. Against 1), BATON makes the subtask the unit of exploration. Each subtask is explored in the inexpensive short-horizon regime, and its solution is stored in memory. A long-horizon trajectory is then composed from these solutions rather than discovered whole. Exploration cost thus becomes linear in the number of stages (), and every failure is attributed to a single stage. Against 2), BATON equips this exploration with a transition-aware memory. Within a subtask, a verifier agent governs the invocation transition: the VLA is invoked only after the wrist view confirms that the scene is ready. Across subtasks, handoff transition restores an entry state disturbed by the predecessor's residue, and lookahead transition determines the execution strategy whose outcome the successor can inherit. On the long-horizon manipulation benchmark RoboMemArena, BATON improves task success rate by 37.7% and cumulative success rate by 29.7% over the current SoTA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.