MosaicAttention: Dual-Pooled Sparse Attention for Accelerating Multi-View 3D Geometry Transformers
Abstract
Feed-forward multi-view geometry Transformers recover cameras and scene geometry from images in a single forward pass by aggregating cross-view evidence through global attention. However, the global attention scales quadratically with the number of input views, making long-sequence inference extremely expensive. Block-sparse attention here provides an intuitive way to accelerate global attention. Existing methods typically treat visual tokens as a flattened 1D sequence and represent each block through only average statistics with uniform sparsity across all layers, which overlooks the 2D spatial layout, salient variations within each block, and layer-wise varying sensitivity to block removal, leading to unreliable block-importance estimation and potential performance loss. These limitations motivate us to introduce MosaicAttention, a sparse attention framework that jointly exploits both 2D average- and maximum-pooled views for more reliable block routing. Our theoretical analysis shows that Avg preserves mean regional logits while Max captures additional within-region variation, motivating dual pooling over 2D visual regions as coarse-grained proxies of token interactions. We then calibrate these dual-pooled proxies against dense value-weighted output contributions to obtain reliable block priorities, and allocate computation according to both block importance and layer-wise sparsification sensitivity. We further bound regional logit approximation and sparse-attention output errors and characterize when calibration preserves block selection. Experiments on multiple feed-forward geometry Transformers and 3D prediction tasks show that MosaicAttention achieves a favorable accuracy-efficiency trade-off over previous methods. Especially, on inputs with 128 views, it achieves up to 2.89 end-to-end inference speedup while maintaining competitive performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.