GeoCircuit: Training-Free Task-Aware Head-Circuit Extraction for Visual Geometry Transformers
Abstract
Feed-forward 3D reconstruction models such as the Visual Geometry Grounded Transformer (VGGT) use a single network to predict camera pose, depth, and point maps. However, their deployment has been constrained by the billion-scale parameter and quadratic complexity. Existing VGGT acceleration methods mainly reduce tokens or sparsify attention entries, while the redundancy and task-dependent roles of complete attention heads remain underexplored. We instead ask whether a frozen visual geometry transformer contains smaller, structurally executable head circuits specialized for different geometric task sets. To investigate this question, we present GeoCircuit, which instruments both frame and global attention with complete-head gates and treats the pretrained model as a frozen parent network of candidate geometric subnetworks. Using dense-model predictions as label-free task teachers, GeoCircuit estimates task-aware gate Gauss-Newton matrices and applies shrinkage to preserve useful cross-head interactions while suppressing estimation noise. Given a target head retention ratio, GeoCircuit directly selects a fixed-budget circuit through a normalized multi-task objective, producing subnets of different sizes without fine-tuning. With 50% of heads retained, GeoCircuit outperforms state-of-the-art method at the corresponding retention setting by 5.93 percentage points on camera pose estimation and improves the point map estimation results by 43.5%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.