acceptodds
Under review as a conference paper at ICLR 2027

GeoPACT: Geometry-Preserving Token Compression for Spatial Reasoning in VLMs

Abstract

Vision-language models excel at semantic recognition and visual question answering, yet remain unreliable at spatial understanding: eg. metric distance, object size, relative direction, and multi-view scene geometry. This limitation becomes critical for embodied agents, where visual context grows continuously while memory and latency budgets remain constrained. Although geometry-aware encoders provide rich spatial representations, their dense patch features are costly to process and poorly aligned with language-model embeddings. Compressing them aggressively is therefore essential—but conventional language-only training can discard precisely the local and metric information needed for spatial reasoning. In this paper, we introduce a general training scheme that explicitly preserves spatial information during compression. During projector pretraining, lightweight auxiliary heads reconstruct dense segmentation and depth targets from the compressed tokens, encouraging them to retain object-level and geometric cues. The heads are removed before downstream fine-tuning, adding no inference cost. We validate this approach with two complementary encoders: DUNE, which provides fine-grained semantic features, and MapAnything, which captures metric multiview geometry. We then introduce GeoPACT, a lightweight dual-encoder VLM that compresses these representations into 20 tokens per view and 200 scene-level tokens—only 840 visual tokens for 32 views—and couples them with a 0.8B language model. Experiments on 2D and 3D spatial-reasoning benchmarks show consistent improvements, while ablations and frozen-token probes confirm that segmentation and depth information remain more accessible after compression. GeoPACT demonstrates that strong visual compression can preserve geometry, enabling scalable spatial reasoning for embodied VLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.