Persistent3D: Open-Vocabulary Persistent Object spatial Perception from Monocular Video via Cross-Frame Instance Association
Abstract
Spatial perception for embodied intelligence requires persistently recognizing and understanding scene instances during dynamic 3D reconstruction. However, cross-frame instance association fails to maintain stable object identities under occlusion, visually similar instances, and substantial viewpoint changes. This prevents consistent aggregation of 3D information across views. To address this issue, we propose Persistent3D, an open-vocabulary framework for persistent object spatial perception from monocular video. Persistent3D first recovers 3D scene geometry and camera poses from video, and combines VLM-based open-vocabulary recognition with instance segmentation to construct per-frame object observations. Starting from frame-level instances, bidirectional temporal propagation and one-to-one mask matching are performed within local temporal windows to alleviate association failures caused by short-term occlusion and viewpoint variation. Meanwhile, a same-frame mutual exclusion constraint is introduced to maintain local cross-frame temporal associations and suppress erroneous merging between visually similar instances. For local identity fragmentation caused by long-term occlusion or viewpoint changes, instance reconnection is performed by leveraging scene-point index overlap, color distributions, and frame-set mutual exclusion constraints across different identity fragments in the 3D scene. During persistent spatial perception, a sequence-wide instance association graph is dynamically constructed to maintain stable persistent identities across frames. Building on this, Persistent3D fuses multi-view vision-language evidence to estimate functional object orientations and fits semantically oriented bounding boxes. The final unified 3D object representation integrates persistent identity, open-vocabulary semantics, 3D geometry, and functional orientation. Extensive experiments demonstrate consistent and effective improvements in cross-frame instance association, 3D information aggregation, and functional orientation estimation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.