TEAM-VLA Agents: Chain-of-Thought Reasoning for Multi-Agent Embodied Autonomy with Closed-Loop Recovery
Abstract
For Vision-Language-Action (VLA) based embodied robots, it is essential to establish an autonomous pipeline in open environments, which spans the entire process from goal-driven planning to execution verification and fault recovery. Existing methods often couple high-level planning with reactive execution, lacking explicit reasoning and closed-loop recovery. To address these limitations, we propose the TEAM-VLA Agents (TVA) framework, which adopts a novel paradigm that integrates explicit Chain-of-Thought (CoT) reasoning with multi-agent team collaboration. By leveraging role-based task decomposition and a context-sharing architecture, TVA constructs a CoT-driven autonomous system that unifies task planning, execution monitoring, as well as autonomous failure diagnosis and recovery within a single closed-loop process. This design not only ensures interpretable and traceable task execution but also effectively mitigates error cascades. Empirical evaluations demonstrate that TVA substantially improves success rates on long-horizon tasks and exhibits enhanced robustness and generalizable recovery capabilities across diverse execution failures. Our findings underscore that explicitly incorporating CoT-based reasoning, task verification, and self-recovery mechanisms is a critical step toward achieving reliable, scalable, and truly autonomous embodied intelligence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.