ELLSA-Bench: A Closed-Loop Benchmark for Omnimodal Human–Scene–Robot Interaction
Abstract
Real-world human–robot interaction is inherently continuous, multimodal and situated: robots must perceive, communicate and act while coordinating with humans and the surrounding environment. Despite rapid progress in embodied models, few support real-time omnimodal interaction, while existing benchmarks lack closed-loop settings that jointly evaluate these capabilities. We introduce ELLSA-Bench, a closed-loop benchmark for naturalistic omnimodal human–scene–robot interaction in simulated environments. ELLSA-Bench explicitly models human behaviors through speech, body motion, gestures and lip movements, and evaluates robots across audio–visual perception, spoken communication and physical action. It comprises 40 atomic scenarios across four families and three interaction levels, covering audiovisual grounding, human–robot collaboration, concurrent communication and action and interruption-aware interaction. To capture the continuous nature of real-world interaction, we incorporate full-duplex tasks requiring robots to listen and observe while speaking and acting. For evaluation, an LLM planner is used to generate scene-grounded interaction plans and synthesize temporally coordinated dialogue, human behaviors and robot actions for scalable naturalistic interaction generation. Finally, we introduce two complementary baselines for full-duplex interaction: a streaming end-to-end audio–visual–action model and an asynchronous agent-based framework, Robo-Harness. Experimental results show that while current models can perform visual perception, speech interaction and action execution, they still struggle with scenarios requiring coordinated interaction across multiple modalities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.