QuadVGGT: Quotient Attention for Accelerating Dense-Token Inference in Visual Geometry Grounded Transformers
Abstract
Feed-forward geometric foundation models such as VGGT have significantly ad- vanced multi-view 3D reconstruction by directly predicting camera parameters and dense scene geometry from uncalibrated images. However, their global at- tention jointly processes all image tokens, causing computational and memory costs to grow rapidly with the number of input views and limiting long-sequence inference. Existing acceleration methods primarily reduce, merge, or compress tokens, potentially altering token-specific interactions within attention. We intro- duce QuadVGGT, a training-free framework based on Quotient Attention. Our key observation is that representational redundancy and computational redun- dancy are distinct: tokens that retain distinct representations for dense geomet- ric prediction can nonetheless induce similar attention interactions, allowing their computations to be shared without merging their representations. QuadVGGT uses sparse operator-induced profiles to select original token representatives for independent key–value and query computation sharing while preserving the orig- inal dense token stream. To further support long-sequence inference, we combine Quotient Attention with scene-adaptive anchor construction and hierarchical fea- ture offloading. QuadVGGT can be directly applied to pretrained VGGT without additional training or modification of its prediction heads. Extensive experiments across dense 3D reconstruction, depth estimation, and camera pose estimation demonstrate substantial efficiency gains while maintaining comparable geometric accuracy. On 1,000-view reconstruction, QuadVGGT achieves a 14.8× speedup and reduces peak GPU memory from 72.4 to 55.9 GB.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.