acceptodds
Under review as a conference paper at ICLR 2027

: Semantic Scale Synthesis for Training-Free Metric Depth from Complementary Foundation Models in Monocular 3DGS SLAM

Abstract

Monocular depth foundation models can predict metric depth from individual images, but these predictions vary in scale and do not provide the multi-view geometric consistency required for estimation of motion while building a map. We introduce -SLAM, a training-free monocular 3D Gaussian Splatting SLAM system that synthesizes two complementary foundation models: UniDepthV2, which infers single-frame metric depth, and VGGT, which infers multi-view-consistent but scale-free geometry. We show that their depth errors are complementary across spatial depth frequencies, introducing an online method of frequency-based fusion that exploits this structure while decoupling the mapping and tracking steps of SLAM. We use persistent semantic object landmarks to further correct frame-to-frame drift on short timescales, and to provide long-term geometric consistency when revisiting previously observed regions. By evaluating this approach across a range of indoor SLAM benchmarks, we demonstrate recovery of consistent metric-scale depth and trajectory estimates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.