MultiView-EB: Benchmarking Ego–Exo Embodied Understanding of MLLMs in Real-World Scenarios
Abstract
Embodied agents rely on egocentric observations for perception and decision-making, yet limited fields of view and frequent occlusions can conceal task-relevant objects, state changes, and spatial relations. Synchronized exocentric views offer complementary evidence, but existing benchmarks focus on single-view reasoning or static cross-view perception. Whether multimodal large language models can integrate both views over time to improve egocentric decisions remains unclear. We introduce MultiView-EB, a benchmark for cross-view spatiotemporal understanding in real-world embodied activities. Built from synchronized recordings of human activities, it contains 2,497 samples across nine activity domains, obtained through semi-automated annotation and human verification. Eight tasks assess cross-view fine-grained perception, task planning, and trajectory prediction through open-ended, multiple-choice, and spatial prediction outputs. We evaluate thirteen general-purpose MLLMs and eight embodied-brain models and analyze performance differences across egocentric and combined-view inputs. On the planning tasks, where input conditions can be matched, adding an exocentric view yields an average gain of only 0.5–2.1 points over egocentric input and degrades accuracy for four of the eight models tested. Additional views thus do not automatically improve embodied reasoning: current models struggle to establish cross-view correspondences, align temporal dynamics, and selectively use complementary evidence for planning. Finally, we introduce MV-EmData-38K, on which fine-tuning improves performance on all eight tasks. MultiView-EB provides a systematic testbed for embodied models capable of reasoning across views.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.