Tapestry: Enabling Continuous Collaboration Across Episodes in Embodied Multi-Agent Systems
Abstract
Embodied multi-agent collaboration is inherently continuous: task execution and intervening environmental changes can reshape the world in which subsequent tasks are performed. Effective collaboration requires agents to accumulate, share, and reuse distributed historical information across episodes, rather than repeatedly starting from scratch. Yet existing embodied multi-agent benchmarks largely isolate episodes, while memory-centric benchmarks primarily study long-horizon historical reasoning within individual embodied tasks, leaving interdependent cross-episode collaboration underexplored. To address this, we introduce the Tapestry Benchmark, to our knowledge the first controlled 3D benchmark for continuous embodied multi-agent collaboration across interdependent episodes. It organizes tasks into continuous streams in evolving environments and provides a unified testbed for studying how embodied agents accumulate, share, and reuse distributed experience under evolving environments and cross-episode continuity dependencies, where earlier task completion, environmental changes, and distributed multi-agent experience jointly determine the context of subsequent tasks. To this end, we propose Tapestry Reasoning, which views the evolving environment as a tapestry woven from distributed observations, actions, state transitions, and execution outcomes across agents and episodes. Tapestry Weaving (TW) interweaves fragmented historical evidence into coherent cross-episode histories, while Tapestry Evidence Grounding (TEG) retrieves and grounds relevant historical evidence to interpret current instructions, infer implicit task-relevant information such as object identities, states, locations, and configurations, and translate these into executable decisions. Evaluations on the benchmark demonstrate that Tapestry Reasoning consistently outperforms existing methods under matched backbone and execution settings across diverse LLM backbones. Extensive experiments further demonstrate its robustness under historical and execution interference and validate the increasing difficulty introduced by the proposed continuity dependencies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.