acceptodds
Under review as a conference paper at ICLR 2027

Ground3R: Streaming 3D Referring Expression Segmentation with Recurrent Multimodal Memory

Abstract

3D referring expression segmentation (3D-RES) aims to segment objects described by natural language in 3D scenes. Existing methods typically rely on preconstructed point clouds or offline collections of RGB images, limiting their applicability to embodied agents that explore unfamiliar scenes through streaming observations. In such scenarios, agents must reconstruct the scene and identify the referred object incrementally as new images arrive, without access to future observations. To address this setting, we introduce streaming 3D referring expression segmentation (Stream3D-RES), which jointly reconstructs the scene and segments the referred object from incoming RGB images using only current and past observations. This task presents two key challenges: (1) identifying the target may require contextual information observed in earlier views, and (2) distinguishing the referent from similar objects becomes difficult as the target leaves and re-enters the field of view. To tackle these challenges, we propose Ground3R, an online framework that combines recurrent multimodal memory with spatially localized object queries. Its Multimodal Memory Branch maintains a language-conditioned scene state across frames, allowing current predictions to leverage historical context. The Spatial Query Localization and Refinement module generates language-independent object proposals and uses the expression and accumulated scene context to identify the referent and refine its mask. Query features are transferred across views through 3D centroid matching, while explicit target-visibility prediction enables empty masks when the target is absent. We further introduce StreamRefer, a benchmark covering four target-visibility patterns and three observation lengths. Extensive experiments demonstrate that Ground3R achieves state-of-the-art performance among online models, substantially outperforming causal 3D baselines and 2D-lifting alternatives across all visibility patterns.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.