acceptodds
Under review as a conference paper at ICLR 2027

EgoWorldMem: 4D World Memory for Long-Horizon Egocentric Video Understanding

Abstract

Egocentric videos contain substantial redundancy as the same physical environment is repeatedly observed from changing viewpoints, while brief object interactions often carry critical evidence. Existing memory mechanisms typically compress visual history based on temporal proximity or feature similarity, making it difficult to distinguish camera motion from actual changes in the world. We introduce **EgoWorldMem**, a 4D world-memory framework that compacts redundant observations of the physical world rather than merely historical tokens. EgoWorldMem reprojects an evolving geometric map into each incoming view to compensate for ego motion, retains sparse world-grounded visual evidence, and organizes dynamic observations into object-centric event segments through persistent tracking. The resulting memory is serialized into a compact visual-token sequence for VLM inference. Across five video-understanding benchmarks, EgoWorldMem improves the average accuracy of Qwen3-VL and LLaVA-NeXT-Video by 10.0 and 23.6 percentage points while removing 73.96% and 78.65% of visual tokens, respectively. An online variant further reduces time-to-first-token by 57.39% on EgoSchema, demonstrating the effectiveness of world-grounded memory for efficient long-horizon egocentric video understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.