GeoRank: Geometry-Aware Low-Rank Compression for Efficient 3D Reconstruction
Abstract
Feed-forward visual geometry transformers have advanced 3D reconstruction, but their large backbones incur substantial computational and memory costs. Existing efficiency methods primarily reduce redundancy in spatial tokens; we instead investigate a complementary and largely overlooked source: channel-space redundancy. We find substantial low-rank structure in both encoder representations and decoder operators, but its geometric importance is non-uniform across tasks and not fully captured by activation energy alone. Motivated by this, we introduce GeoRank, a geometry-aware low-rank compression framework for feed-forward 3D reconstruction. In the calibration stage, GeoRank first conducts activation-aware factorization to construct low-rank operators that better preserve the model’s effective feature space. Since activation energy does not fully capture geometric importance, we further redistribute a rank budget according to geometric sensitivity. During training, to handle variation beyond static compression, we introduce lightweight adaptive refinement that adjusts the compressed representation to each input and selectively restores information lost during truncation. Geometry-sensitive regularization further stabilizes geometry-critical directions. Furthermore, the factorized representation naturally extends to streaming inference by caching historical features in compact rank space. Benchmarking results demonstrate that GeoRank substantially reduces model cost while retaining strong reconstruction quality. Compared with the full model, it reduces decoder parameter size by 45.7% and GFLOPs by 31.0%, while improving 7-Scenes point-map accuracy from 0.055 to 0.035 and completion from 0.090 to 0.050. The same compact representation also extends effectively to causal reconstruction with strong online performance. Overall, GeoRank establishes channel-space low-rank compression as a complementary direction for building efficient feed-forward and streaming 3D reconstruction models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.