acceptodds
Under review as a conference paper at ICLR 2027

GeoCoCo: Representation-Decoupled Co-occurrence Compression for MLLM Geospatial Tokens

Abstract

Long visual token sequences (e.g., -) from ultra-high-resolution (UHR) remote sensing imagery increase the inference cost of multimodal large language models (MLLMs). Existing token compression methods typically depend on high-dimensional MLLM features, adding planning overhead and requiring feature computation before planning. Our observation study shows that low-dimensional image descriptors provide compression guidance with comparable accuracy and lower prefill time, while spatially regional grouping further improves accuracy. We therefore propose GeoCoCo, a feature-decoupled, plug-and-play framework for low-overhead, region-level visual token merging. Offline, GeoCoCo quantizes image descriptors into shared basic visual symbols and hierarchically merges them using spatial adjacency co-occurrence statistics aggregated across images to learn a shared compression prior. Online, this prior guides connected-region planning from input images and aggregation of the original model's visual token features without retraining the target MLLM. Planning is independent of MLLM features, enabling parallel CPU planning and GPU visual encoding. Experiments across XLRS-Bench-lite, LRS-VQA, and MME-RealWorld-RS and six MLLMs show that GeoCoCo reduces prefill time by up to relative to uncompressed models, achieving the best performance–efficiency trade-off among the compared compression methods. We will release the relevant code for further research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.