: Semantic Scale Synthesis for Training-Free Metric Depth from Complementary Foundation Models in Monocular 3DGS SLAM
Abstract
Monocular depth foundation models can predict metric depth from individual images, but these predictions vary in scale and do not provide the multi-view geometric consistency required for estimation of motion while building a map. We introduce -SLAM, a training-free monocular 3D Gaussian Splatting SLAM system that synthesizes two complementary foundation models: UniDepthV2, which infers single-frame metric depth, and VGGT, which infers multi-view-consistent but scale-free geometry. We show that their depth errors are complementary across spatial depth frequencies, introducing an online method of frequency-based fusion that exploits this structure while decoupling the mapping and tracking steps of SLAM. We use persistent semantic object landmarks to further correct frame-to-frame drift on short timescales, and to provide long-term geometric consistency when revisiting previously observed regions. By evaluating this approach across a range of indoor SLAM benchmarks, we demonstrate recovery of consistent metric-scale depth and trajectory estimates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.