Geometrize What Really Matters: Toward Lighter and Faster Spatial MLLMs
Abstract
Spatial multimodal large language models augment vision-language reasoning with explicit geometric representations, enabling understanding of object locations, distances, directions, and cross-view relations. This capability, however, introduces redundancy at two levels: dense multi-view observations must first pass through expensive geometric encoders, and the resulting geometry-enhanced visual tokens then increase language-model prefill computation and KV-cache storage. Existing pruning methods typically optimize only the latter stage after geometry has already been computed, while geometric diversity alone may fail to preserve evidence that is critical to the frozen model. We present GRAFT (Geometry Routing Across Frames and Tokens), a training-free hierarchical routing framework that allocates geometric computation and evidence along the inference pipeline. Before full geometric encoding, GRAFT uses shallow features from a frozen geometry backbone to select views with reliable cross-view relations and temporal coverage, reducing the observations processed by the full geometry encoder. After visual-geometric encoding, it estimates model sensitivity to construct a compact semantic evidence set and exploits aligned 3D coordinates to correct spatially invalid token substitutions under an exact budget. Extensive experiments on three benchmarks demonstrate that GRAFT consistently maintains competitive spatial reasoning accuracy while reducing both geometric encoding and language inference costs. On VSI-Bench, retaining only 5% of visual tokens preserves 94% of the accuracy, while reducing inference latency by 82% and peak memory usage by 55%. Anonymously released code: https://anonymous.4open.science/r/graft-88DA
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.