acceptodds
Under review as a conference paper at ICLR 2027

CalibAdapter: Adapting 3D Foundation Models to Calibrated Cameras

Abstract

Recent feed-forward 3D foundation models have demonstrated strong capabilities in reconstructing scene geometry, supporting a wide range of downstream tasks. In practical settings such as robotic manipulation and autonomous driving, calibrated camera intrinsics and extrinsics are usually readily available through standard sensor-calibration pipelines. Existing works like MapAnything, DepthAnything3, and PromptDA have already shown that utilizing calibrated camera parameters can improve reconstruction accuracy. However, these methods do not guarantee that the predicted depth will back-project to geometry consistent with the provided calibration. We quantify this deployment discrepancy using **Calibrated Pointmap Error (CPE)**, which measures the resulting pixel-corresponded metric 3D error. To address this discrepancy, we propose **CalibAdapter**, a lightweight adaptation mechanism for incorporating known camera calibration into 3D foundation models. CalibAdapter attaches a projective sparse residual branch to a selected cross-view interaction block while keeping all pretrained parameters frozen. CalibAdapter consistently reduces CPE across **VGGT-Ω**, **VGGT**, **DA3**, and **π³**, achieving an average relative CPE reduction of **28.3%** together with overall improvements in conventional reconstruction metrics. It further yields gains on several held-out benchmarks. Beyond conventional reconstruction benchmarks, CalibAdapter improves geometric accuracy in **autonomous driving with known rig calibration and ego poses**, and increases **real-world robotic manipulation** success rate by 18.4 percentage points over the parameter-matched baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.