acceptodds
Under review as a conference paper at ICLR 2027

VLVLM: Learning Variable Length Visual Tokens for Efficient Vision Language Models

Abstract

Image tokens account for much of the compute cost of vision language models (VLMs), particularly at high resolution. Prior research has sought to reduce this cost by pruning tokens deemed redundant or uninformative, primarily through training-free approaches focusing on general visual understanding tasks. In this work, we show that these approaches incur substantial accuracy losses on text-rich images, where they are outperformed by a simple finetuning strategy that pools consecutive fixed-size groups of image tokens. However, fixed pooling is fundamentally limited in that it compresses all regions uniformly, despite differences in the amount of information they might contain. We introduce VLVLM, a training-based approach that dynamically chunks image tokens using a lightweight boundary predictor. At 4× compression (i.e., using only 25% of the original image tokens), VLVLM retains 89% of the uncompressed models’ performance on text-heavy visual benchmarks and outperforms both fixed pooling (87%) and training-free compression (82%), while preserving 99% performance on general tasks. Notably, the advantage persists under a more aggressive 8×compression: we obtain 85% retention compared with 81% for fixed pooling. These gains hold across VLM families, image encoders, and language model sizes (up to 14B). Further analyses show that the learned boundaries are semantically meaningful, aligning closely with text-rich regions. We will release all our code and checkpoints upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.