A Hierarchical Agent for Long-Horizon Multi-Robot Collaboration with Feedback-Driven Task Continuation
Abstract
Long-horizon multi-robot collaboration requires coordinating operations whose feasibility depends on the outcomes of other robots’ actions. A central challenge is aligning high-level decisions with physical execution: an unmet prerequisite, a failed operation, or a changed scene state can invalidate a planned next step. We present a hierarchical agent that coordinates multiple robots through a common execution interface. A vision–language agent selects robot roles and skills from the task, observations, structured physical-state feedback, capabilities, and execution history. The interface validates requests, dispatches learned manipulation and supporting skills, and returns outcomes or unmet conditions for the next decision. It also supports replaceable coordination models and skill backends. We evaluate direct handover, object-dependent sorting, model configurations, and task continuation in a four-robot simulation. Across 75 sorting episodes, the agent selects every receiver and destination correctly and completes 62 deliveries (82.7%). Under 45 recoverable post-handover drops, feedback-driven coordination completes 15 deliveries, whereas a one-shot plan with the same execution components completes none. A paired 2 × 2 study shows both high-level models operating with both backends and selecting all destinations correctly, while delivery varies with low-level performance. The results show that aligning decisions with physical feedback supports continuous collaboration across robots and configurable components.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.