acceptodds
Under review as a conference paper at ICLR 2027

See Once, Ask Many: Object-Anchored Neural Memory for Multi-View Spatial Reasoning

Abstract

Multi-view spatial reasoning requires consistent object identity and reference-frame reasoning across partial observations, which remain difficult for vision-language models (VLMs). Existing methods often use external 3D models to provide spatial features, which VLMs then combine with visual semantics to directly predict spatial answers. We introduce See Once, Ask Many (SOAM), which learns object-addressable neural scene memory within the VLM and uses it to compose question-specific programs for explicit geometric execution. Given sparse views, a single VLM performs neural scene compilation: it organizes cross-view observations into anonymous objects and binds their visual representations into a question-independent neural memory. Object identities provide shared addresses for semantic grounding, view-local localization, and geometric computation. The same VLM then interprets each question through this memory and composes a typed program, while a frozen geometric backend supplies measurements for execution. The resulting scene representation supports subsequent questions. Joint state–program supervision and reuse-aware policy optimization train the shared memory for observation fidelity and utility across questions, with execution feedback assigned to each program individually. SOAM-8B outperforms the strongest baseline on VIEW2SPACE-v1 by 6.96 points in grounding mIoU and 5.28 percentage points in MCQ accuracy, while achieving a macro score of 72.85 on SPAR-Bench. Ablations show higher task scores with object-anchored memory and full training than with their respective controls.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.