EgoMM: Towards Multimodal Egocentric Intelligence
Abstract
Wearables capture synchronized video, audio, and inertial signals, yet existing egocentric understanding systems remain largely centered on vision. When cameras are occluded, poorly lit, or directed away from an event, complementary evidence can remain in sound and motion. Learning from this evidence requires models that integrate wearable signals and evaluations that distinguish their contributions. We present EgoMM, connecting benchmarking, multimodal learning, and long-horizon memory around this goal. We establish EgoMMU, to our knowledge the first benchmark to assess egocentric language models across visual, acoustic, and inertial tasks with explicit modality requirements. Its five task families form thirteen task-modality combinations, supported by construction-time checks and controlled ablations. We then develop EgoMind, which connects all three modalities to a vision-language backbone through staged sensor pretraining and modality-matched instruction tuning. At each of four model sizes, a single checkpoint accepts any non-empty subset of the streams. Pretraining improves all five task families at every size, alongside competitive performance across external egocentric benchmarks. Finally, we build EgoBuddy to organize multimodal observations into hierarchical text memory for long-horizon egocentric understanding. Video-audio evaluations examine retrieval and question answering under progressively smaller visual frame budgets. Together, these components advance egocentric understanding toward a richer account of what is "seen" and "heard", and how the wearers "move".
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.