Token-Adaptive Routing of Multi-Level Geometry for Spatial Reasoning
Abstract
Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearance. Recent multimodal models incorporate geometric information through structural branches, 3D-aware supervision, reasoning-stage fusion, or long-horizon memory. While these approaches demonstrate the importance of geometry for spatial intelligence, many combine geometric features with visual representations through shared or stage-wise rules. This leaves a finer-grained question: which levels of geometric evidence should each visual token use for its spatial context? To address this question, we introduce GeoWeaver, a token-adaptive geometric grounding framework that selects multi-level evidence before language decoding. GeoWeaver constructs a multi-level geometry bank from a frozen geometry encoder, and each visual token sparsely selects evidence from this bank. The selected evidence is incorporated into visual tokens via a residual grounding operation prior to language modeling, yielding geometry-grounded representations for downstream reasoning. Across six reported spatial reasoning benchmarks, GeoWeaver attains a 64.8 mean. In a matched 32-frame ReVSI comparison, it exceeds SpatialStack by 2.9 points, while its four-task general multimodal mean improves on the base model (68.98 versus 68.41). On VSI-Bench, token-adaptive allocation reaches 72.7, compared with 71.30 for global layer weighting and 70.44 for uniform averaging, supporting the value of selecting geometric evidence at token granularity. Code and models will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.