acceptodds
Under review as a conference paper at ICLR 2027

GLINT-3D: Geometry-aware Language-guided Interleaved Compression

Abstract

Spatial understanding from multi-view inputs requires vision-language models (VLMs) to preserve both fine-grained visual evidence and geometric relationships that can only be established across views. Feeding per-frame semantic and geometry-aware features directly to the language model is effective but computationally expensive, since the visual token sequence, and with it the decoder KV cache, grows linearly with the number of input frames. In this work we introduce GLINT-3D, a multi-view VLM built on language-guided, interleaved compression of semantic and geometric features. GLINT-3D fuses per-frame features from the VLM vision encoder with multi-view 3D-aware features from a pretrained 3D foundation model, then compresses the result into scene-level and frame-level latent tokens. Rather than compressing once before the language model, GLINT-3D initializes these latent tokens from learnable embeddings and progressively injects visual information through compression layers interleaved with the VLM layers, allowing the question tokens to guide what information is retained. Experiments on a suite of spatial and general visual benchmarks show that one-shot compression causes substantial information loss, while the combination of interleaving and language guidance largely closes the gap, approaching the mean score of the uncompressed model while substantially reducing the visual token count and decoder KV-cache size. GLINT-3D thus offers a favorable trade-off between spatial reasoning performance and scalability to larger image sets and language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.