SpatialMem: Object-Level 3D Memory Boosts Spatial Understanding from Monocular Videos
Abstract
Spatial understanding from monocular video requires locating objects in 3D and recognizing the same object across different views. Predicted depth and camera poses place observations in a common 3D coordinate system, but connected regions in the reconstruction do not always correspond to individual objects. Reconstruction artifacts may join geometry from neighboring objects, while occlusion may leave a single object represented by disconnected fragments. These ambiguities can compromise the semantic and geometric consistency of scene representations, undermining spatial understanding. To this end, we introduce SpatialMem, a framework that constructs reusable object-level 3D memory from monocular video by combining predicted geometry with cross-view visual evidence. The system associates semantic masks across views through the reconstructed geometry and uses visual evidence to determine which observations belong to the same object. It stores each instance with its observed geometry and links to supporting views. The memory supports both read-only spatial programs and a lightweight latent-injection interface to a frozen vision-language model. Using up to 300 sampled RGB frames per video for memory construction, SpatialMem achieves a seven-task average of 65.3% on ReVSI, the highest among the compared methods. In exploratory evaluations, latent features extracted from this memory also improve exact-match accuracy from 53.2% to 58.2% on SQA3D and from 28.6% to 30.6% on ScanQA. Our code will be released at https://anonymous.4open.science/r/SpatialMem-7E3D/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.