acceptodds
Under review as a conference paper at ICLR 2027

Visual Grounding and Deliberation Help Multi-Robot Teams Make High-Stakes Decisions

Abstract

Humans have developed communication skills to disambiguate and coordinate during high-stakes collaborative tasks. Whether teams of robotic systems can do the same remains unclear, especially when no single robot has full access to the information necessary to make the correct decision. With recent developments in Vision-Language Model (VLM)-powered generalist robotic systems, such scenarios are becoming prevalent and have raised serious safety concerns. To address the gap in understanding such systems, we study this problem with a benchmark across three household environments modeled as Dec-POMDPs with communication, in which a team of robots must jointly decide whether a proposed action should Proceed or Hold. We find that unstructured communication is unreliable: team accuracy can fall to chance level, and common changes to prompting, communication topology, team size, and interaction rounds do not consistently resolve these failures. Further, an idealized full-information diagnostic, in which every robot receives all observations and local context, raises accuracy to 100% in some settings, indicating that handling distributed evidence is a major source of error. To address this, we introduce a structured deliberation protocol, called HUDDLE, that requires robots to explicitly share local context, ground visual claims in observations, and assess the pooled evidence before deciding. HUDDLE substantially improves decision accuracy, reaching up to 96% in our best setting. These findings highlight communication skills as the key bottleneck for reliable multi-robot decision-making and provide a foundation for designing safer, more collaborative autonomous systems in partially observable environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.