Extending 3D Foundation Models from to Any Central Camera with Distortion Awareness Token
Abstract
3D foundation models, such as Visual Geometry Grounded Transformer (VGGT), have shown superior performance in inferring 3D attributes of a scene from several views. However, they do not generalize to non-perspective cameras, limiting their usage in other camera models, such as fisheye and panoramic cameras. To this end, we propose VGGT-Any, which extends VGGT to any type of central camera and to heterogeneous camera systems without retraining the base model. At the core of VGGT-Any is a novel camera representation: it decomposes camera geometry into a canonical space, which follows the usual perspective-camera formulation, and a distortion space, which expresses each camera's specific projection. We further introduce a distortion token, which is injected into the frozen foundation model to reconstruct scenes from various central cameras, or even from heterogeneous camera systems. Moreover, since existing datasets each cover only a limited range of camera types, we construct LensZoo, a synthetic indoor dataset in which any central camera model can be synthesized at shared poses, enabling controlled comparisons across camera models. Our evaluations show that VGGT-Any is competitive with camera-specific and fine-tuned baselines on real fisheye and panoramic benchmarks, and exactly retains the base model's pinhole reconstruction ability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.