acceptodds
Under review as a conference paper at ICLR 2027

EgoMM: Towards Multimodal Egocentric Intelligence

Abstract

Wearables capture synchronized video, audio, and inertial signals, yet existing egocentric understanding systems remain largely centered on vision. When cameras are occluded, poorly lit, or directed away from an event, complementary evidence can remain in sound and motion. Learning from this evidence requires models that integrate wearable signals and evaluations that distinguish their contributions. We present EgoMM, connecting benchmarking, multimodal learning, and long-horizon memory around this goal. We establish EgoMMU, to our knowledge the first benchmark to assess egocentric language models across visual, acoustic, and inertial tasks with explicit modality requirements. Its five task families form thirteen task-modality combinations, supported by construction-time checks and controlled ablations. We then develop EgoMind, which connects all three modalities to a vision-language backbone through staged sensor pretraining and modality-matched instruction tuning. At each of four model sizes, a single checkpoint accepts any non-empty subset of the streams. Pretraining improves all five task families at every size, alongside competitive performance across external egocentric benchmarks. Finally, we build EgoBuddy to organize multimodal observations into hierarchical text memory for long-horizon egocentric understanding. Video-audio evaluations examine retrieval and question answering under progressively smaller visual frame budgets. Together, these components advance egocentric understanding toward a richer account of what is "seen" and "heard", and how the wearers "move".

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.