Let the Mind Conceive, Let the Hand Achieve: Following Multimodal Instructions for Long-Horizon Robotic Manipulation
Abstract
Following a complete robot instruction requires maintaining task continuity and preserving the connection between intent and physical referents as the scene changes. These challenges are coupled: executing the correct sequence is insufficient if actions target the wrong entities, and grounding the correct entities is insufficient if operations are repeated or omitted. We introduce Orchestra, an agent harness that separates task-level interpretation and execution management from local action generation. Given a complete multimodal instruction, Orchestra constructs structured subtasks and maintains a shared context aligning the active goal, spatial grounding, observation history, and task progress. A spatially and temporally grounded action policy and an independent progress evaluator use aligned context to generate actions and guide subtask transitions. We also introduce RMBench-Compose to evaluate longer compositions of skills learned from atomic instructions and short trajectories. Orchestra achieves mean success across nine RMBench tasks, exceeding Mem-0 by percentage points. On RMBench-Compose, it achieves an Average Success Length of subtasks, times that of HarnessVLA. Across four real-world tasks, it improves mean Average Success Length by over its base-policy variant. These results support shared-context coordination as an effective approach to sustaining grounded instruction following over long horizons.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.