acceptodds
Under review as a conference paper at ICLR 2027

Geometry-Aware Memory Coordination for Spatial Understanding from Long Videos

Abstract

Spatial understanding from long videos requires integrating partial observations across viewpoints and recalling past content within a limited memory budget. Repeated visits can reveal complementary views or changes in familiar regions. Existing memory methods use predictive novelty to guide retention or enrich language inputs with geometric features. Yet novelty under viewpoint changes need not reflect distinct content, and an observation can be redundant for geometry yet retain unique content for future questions. We introduce GEOMEM, a geometry-aware framework that coordinates visuospatial and episodic memories. Visuospatial memory provides historical context for geometry encoding of new views. The resulting features enrich the visual inputs used to form episodic memory, which retains observed content and its temporal context for later questions. Shared cross-view evidence guides separate retention decisions for geometric support and content under fixed budgets. Historical content can therefore remain accessible even when its source view becomes geometrically replaceable. Experiments on VSI-Bench and VSI-SUPER show leading performance in spatial understanding, long-horizon recall, and counting. Budget analyses show improved geometry estimation and question answering at matched memory budgets. Ablations support geometry fusion and complementary memory updates. Combining geometric association with content comparison outperforms either criterion alone. Streaming measurements show stable peak GPU memory and per-frame observation time on a four-hour video. Code and scripts will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.