acceptodds
Under review as a conference paper at ICLR 2027

3DMem: Synergizing Textual and Spatial Memories for Conversational 3D Scene Understanding

Abstract

Existing 3D vision-language models are primarily developed for single-turn interactions, whereas natural dialogue requires a model to retain references and incorporate new constraints across turns. We introduce Conversational 3D Scene Understanding and ScanFlow, a benchmark for multi-turn question answering and grounding. Our framework, 3DMem, maintains two complementary recurrent states: a cumulative textual summary and a continuous spatial field aligned with scene superpoints. The spatial state is updated from the current causal prefix and previous predicted region, then read through language-side embeddings and residual modulation of grounding features. We construct continuous spatial supervision directly from point-cloud geometry, separating memory-region targets from exact instance-mask targets. Two supervised stages learn memory prediction and task responses using predicted memory trajectories, followed by IoU-aware optimization of the textual-memory policy with supervised grounding updates. This formulation connects evolving dialogue context to an explicit scene-indexed representation without generating discrete spatial buckets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.