EmbodiedScope: Benchmarking Agentic Capabilities in Interactive Embodied Tasks
Abstract
Recent advances in general-purpose multimodal large language models (MLLMs) are converting them from passive perception systems into embodied AI agents that can perceive, reason, and act through interaction. Reliable execution of complex tasks requires agents to acquire missing evidence, retain task-relevant history, reason about physical constraints and action consequences, and respond to execution feedback. To assess the capabilities of embodied AI agents, we introduce EmbodiedScope, a diagnostic benchmark built on BEHAVIOR with 240 closed-loop household tasks under a humanoid setting. Embodied AI agents receive egocentric visual observations and act through a shared intermediate interface that abstracts fine-grained motor control while preserving task-level decisions. Controlled variants manipulate task-relevant evidence, interaction history, physical conditions, and execution outcomes, while tasks span basic interaction, capability-focused tasks, and structured multi-capability tasks. We evaluate representative MLLMs, including paired base and harnessed configurations for selected backbones. The best-performing system is comparatively stronger on Active Information Acquisition and Physical Reasoning, but weaker on Memory and Verification and Recovery. Performance drops sharply on L2, where multiple demands must persist and interact within the same episode. Agent-system support improves performance on selected capability subsets, but only partially mitigates the degradation on structured multi-capability tasks. EmbodiedScope bridges capability-level diagnosis and integrated embodied execution, enabling systematic studies of how capabilities that appear reliable in focused settings persist, interact, and fail as task demands become increasingly integrated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.