GeoLens: Where to Prune Transfers Across 3D Foundation Models, How Much Does Not
Abstract
Pruning is a standard way to make large pretrained models cheaper to run, and 3D foundation models need it, above all in robotics and driving. They have become the default front-end for geometry from images, returning depth, point maps and camera pose in one feedforward pass, and on the multiview, streaming and recurrent models the bypassable blocks take a quarter to half of the latency. Yet in a model that emits dense geometry a badly chosen block warps the scene, and to our knowledge whether training-free block pruning is safe there has not been evaluated. GeoLens is the first cross-architecture evaluation of it: eight architectures across four inference paradigms and 26 indoor RGB-D sequences, 208 cells under one depth acceptance test, plus 304 cells on ETH3D, TartanAir and ScanNet, every number regenerated from released per-cell files. It yields a recipe and two limits. Where to prune transfers in domain: one fixed set of five blocks per architecture, taken from unlabeled scoring of its other sequences, passes 94% of cells and matches per-sequence re-scoring cell for cell. An oracle that spent a labeled pass on every block, measuring its damage on the benchmark itself, gains at most about two points on that set (95% interval): one unlabeled choice per architecture suffices, with no retraining. A prior fit on other architectures picks sets that pass 87% of cells in its seven-model pool and 17/26 outside it. How much is the first limit: capacity, the largest bypass ratio sustaining ≥80% acceptance, spans 3.2× across the roster (2.0× under a wider block count) and shifts with the evaluation distribution; out of domain no K=5 set we tried clears the bar on three of eight architectures on ScanNet. It does not follow from the per-block ranking, because single-block safety does not compose into set safety, and no pre-registered structural scalar of six tracks it under every block count. The second limit is the one that matters in deployment: the acceptance test licensing the recipe does not certify the task. On nine cells the acceptance test accepts, trajectory error grows by up to 2.54×, across five of the seven architectures we checked. On the multiview, streaming and recurrent models, bypass saves 5.8–11.5% of wall-clock at K=5 and up to 22.5% at capacity. We release the benchmark, the per-cell results and the pre-registrations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.