AdaVGGT: Self-Supervised Test-Time Adaptation for Video-based 3D Reconstruction
Abstract
Feed-forward geometry models reconstruct 3D scenes in a single pass, but repeatedly applying frozen weights to long videos leaves scene-specific errors uncorrected. Conversely, traditional per-scene optimization aligns observations to a target scene but lacks general priors and incurs heavy computation. We bridge these two paradigms with AdaVGGT, a self-supervised online test-time adaptation method that specializes pretrained predictors to a test video on the fly. We freeze the backbone of VGGT and jointly optimize lightweight low-rank (LoRA) corrections on its depth and camera pose heads during inference. Our objective requires no offline training, auxiliary datasets, or external optical flow; instead, it refines predictions directly on current observations via multi-view photometric reprojection and cross-frame depth consistency. To process long sequences online, adaptation proceeds causal-sequentially on overlapping temporal chunks, where each chunk inherits the optimized parameters and optimizer momentum from its predecessor. Across five real-world and synthetic outdoor, driving, and indoor video benchmarks, AdaVGGT consistently improves camera tracking and depth estimation over the unadapted foundation model. Notably, on KITTI Odometry, adaptation cuts the mean tracking error of VGGT from to — a reduction — outperforming both classical vSLAM baselines and existing test-time optimization methods. These results demonstrate how strong general priors and scene-specific self-supervision can work together to enable robust, pixel-aligned 3D reconstruction at deployment. Code and models will be publicly released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.