acceptodds
Under review as a conference paper at ICLR 2027

VASC: Value-Aware Sparse Attention with Cross-layer Memory for Efficient 3D Reconstruction

Abstract

Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pass. However, their quadratic global attention makes long image sequences expensive, while existing sparse methods may favor highly attended yet value-redundant regions. To address these limitations, we introduce VASC, a training-free sparse attention method combining value-aware block selection and execution-aware cross-layer memory. Our value-aware block selection integrates pooled query–key relevance with neighboring value contrast, reducing redundancy while preserving query-relevant and distinctive content. Cross-layer memory tracks unserved demand across layers and updates this state according to actual execution, enabling previously underserved blocks to compete under a fixed computation budget. Experiments on 7Scenes and NeuralRGB-D with VGGT and demonstrate improved pose estimation and reconstruction quality compared with FasterVGGT, together with up to faster inference than dense VGGT. Code and evaluation scripts will be made available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.