PruneBA: Efficient Visual Geometry Initialization with Chunked Progressive Bundle Adjustment
Abstract
Feed-forward visual geometry foundation models provide camera hypotheses for image collections that challenge feature-based Structure-from-Motion (SfM). Turning these hypotheses into supported sparse geometry requires efficient neural inference and effective geometric refinement. We introduce PruneBA, a training-free framework combining shared random patch-token pruning in the Visual Geometry Grounded Transformer (VGGT) with chunked progressive bundle adjustment (BA). Independently inferred overlapping windows undergo local triangulation and BA; a deterministic centrality criterion selects camera estimates for global re-triangulation and BA. Across seven ground–UAV, wide-baseline, and indoor evaluation cases, unpruned progressive BA retains all 157 case-level RGB cameras, compared with 58 for classical incremental SfM, and lowers internal mean reprojection error relative to direct global BA in every case. On Bus Ground–UAV 30, progressive refinement increases point yield from 96 to 298; adding 75% pruning yields 364 points at 1.364 pixels while retaining all 30 cameras. Synthetic ten-frame inference achieves a 2.63× speedup with random weights and 2.20× with pretrained weights, with lower GPU allocation in both studies. Exact attention-pair counts and finite-population sampling identities characterize the computational and spatial effects of pruning, while a centrality surrogate formalizes camera selection. These results establish sparse neural aggregation and staged geometric refinement as a practical combination for visual geometry initialization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.