Any3D-SLAM: Taming Any General 3D Priors for Robust In-the-Wild SLAM
Abstract
We introduce Any3D-SLAM, a robust visual SLAM system for jointly recovering camera poses and dense depth from streaming videos under unconstrained capture conditions. Conventional visual SLAM achieves accurate camera estimation by explicitly modeling multi-view geometric constraints, but its reliance on reliable correspondences makes it vulnerable to weak visual evidence and scene dynamics. Existing methods address these challenges with task-specific predictors, but their separate representations limit the use of complementary cues across tasks to support pose estimation. Recent 3D foundation models offer shared multi-frame representations and complementary priors, yet their feed-forward predictions may not fully satisfy multi-view geometric constraints. To harness these representations while enforcing geometric consistency, Any3D-SLAM integrates foundation-model features and complementary depth, motion, pose, and intrinsics priors into an end-to-end differential bundle adjustment (DBA) framework. We develop a warping-based updater for efficient correspondence estimation, substantially reducing memory overhead and enabling richer multi-frame representations and their associated priors to participate in online optimization. Among these priors, we introduce uncertainty-gated relative-pose constraints into online BA to support camera estimation when translational observability is weak or moving objects occlude the static background. Extensive experiments across six benchmarks spanning indoor, outdoor, dynamic, and egocentric environments demonstrate strong camera pose estimation performance in both calibrated and uncalibrated settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.