EgoOmniWorld: Unifying Egocentric Understanding via Ego–Action–World Interactions
Abstract
Egocentric video understanding has rapidly expanded into a diverse landscape of benchmarks targeting different capabilities, modalities, and evaluation formats. While each direction captures an important aspect of egocentric understanding, these capabilities are often studied in isolation, despite being tightly coupled in real-world interaction. We argue that a defining property of egocentric video provides a natural unifying principle: the wearer perceives the world, forms goals, takes actions, changes the world, and observes the resulting feedback to guide subsequent behavior, forming a closed Ego–Action–World interaction loop. Based on this view, we introduce EgoOmniWorld, a unified benchmark that organizes egocentric understanding through four complementary interaction structures: Node, capturing ego, action, and world states; Edge, capturing directed relations among them; Chain, capturing linked dependencies across multiple relations; and Loop, capturing feedback-closed interactions. We further develop an interaction-first, evidence-grounded data pipeline that extracts structured interactions from egocentric videos, grounds them in visual and auditory evidence, and generates questions according to their interaction structure. Rather than treating multimodality and dialogue form as separate task families, the pipeline naturally determines the required modality and question-answering format according to the evidence and dependencies involved. EgoOmniWorld contains 24,166 QA groups, including 20,284 training and 3,882 test samples, spanning diverse interaction scopes, modality requirements, and single- and multi-turn settings. Experiments on representative open- and closed-source multimodal large language models (MLLMs) reveal substantial and uneven limitations across Ego–Action–World interactions. Models trained on EgoOmniWorld consistently improve in-domain interaction understanding and generalize to existing egocentric benchmarks. We hope EgoOmniWorld provides a unified foundation for systematically evaluating and advancing egocentric video understanding. Data, code, and models will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.