acceptodds
Under review as a conference paper at ICLR 2027

TraDE: Trading Cluster Distortion for Efficient Visual Geometry Transformers

Abstract

Visual geometry foundation models enable powerful feed-forward 3D reconstruction, but their reliance on global attention leads to rapidly growing computational costs for long input sequences. Existing training-free token reduction methods alleviate this cost, yet typically determine token selection and retention through predefined regular structures, overlooking the heterogeneous redundancy of token representations. We observe that token similarity is highly non-uniform across input sequences and spatial regions and dynamically evolves with network depth. Based on this observation, we propose TraDE, a training-free framework that adaptively determines both which tokens to retain and how many are needed according to their underlying representations. TraDE formulates token reduction as iterative clustering and uses cluster distortion to balance representation fidelity against the number of retained tokens, while stage-wise repartitioning accommodates evolving representations across depth. Extensive experiments across multiple visual geometry foundation models, datasets, and geometric tasks demonstrate favorable performance–efficiency trade-offs, achieving up to inference acceleration on long sequences. Code is provided in the Supplementary Material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.