acceptodds
Under review as a conference paper at ICLR 2027

Hybrid-VGGT: Scaling Visual Geometry Grounded Transformers to High Resolution via Multi-Scale Detail Injection

Abstract

Feed-forward foundation models have significantly accelerated uncalibrated 3D reconstruction by inferring geometry in a single pass. However, their reliance on global self-attention introduces a severe quadratic computational bottleneck at high resolutions, forcing a compromise between inference efficiency and the preservation of fine-grained geometric details. To address this accuracy-efficiency trade-off, we introduce Hybrid-VGGT, an asymmetric dual-branch extension of VGGT that allocates global multi-view reasoning and high-resolution detail extraction to complementary streams. A low-resolution branch efficiently anchors global scene context and camera prediction, while a parallel Hybrid Detail Extractor (HDE) captures per-view, fine-grained structural cues directly from high-resolution inputs. By integrating multi-layer local convolutions with scalable linear attention, the HDE extracts structural cues across multiple scales. These multi-scale features are then progressively injected into the global low-resolution stream via a customized hierarchical fusion head, effectively refining the coarse geometric representation with high-frequency boundary cues. Comprehensive experiments demonstrate that Hybrid-VGGT achieves state-of-the-art performance on multiple benchmarks for camera pose recovery, dense depth, and point map estimation, yielding high-fidelity 3D geometry and the best overall mean rank in our comparisons. On 100-frame 1036p clips, Hybrid-VGGT runs at 38.56 FPS, compared with 3.12 FPS for VGGT and 5.19 FPS for Pi3.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.