EmboScope: A Multi-Source Comprehensive Benchmark for Embodied Perception, Understanding, and Planning
Abstract
Vision-language models (VLMs) show promise for embodied tasks, yet understanding actions, state changes, and goals remains challenging. Existing benchmarks often struggle to combine broad interaction coverage, fine-grained capability diagnosis, and scalable data construction. We introduce **EmboScope**, a multi-source embodied benchmark with a reusable construction framework. It decouples interaction annotation from rule-based QA generation through a shared schema of objects, actions, and state changes, enforcing task-specific observation constraints while enabling reuse of annotations and task rules. EmboScope integrates human egocentric, real-robot, and simulated interactions across indoor and outdoor settings in everyday and industrial domains. Its 4,467 questions cover 14 tasks at three levels: perception, embodied semantics, and planning and reasoning. We evaluate 23 general-purpose and embodied VLMs, with four additional thinking configurations. The results reveal a gap between local recognition and coherent interaction understanding, alongside complementary strengths across models. Our findings suggest that scaling textual chain-of-thought alone may be insufficient for substantial gains on embodied tasks, motivating stronger visual grounding and interaction understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.