acceptodds
Under review as a conference paper at ICLR 2027

ResizeBench: Revealing the Limits of Resizing for Visual Token Compression

Abstract

Vision-token compression methods are typically evaluated against one another on standard VQA benchmarks, without comparison to pixel-level resizing. Yet on modern variable-resolution VLMs, naive image resizing matches or outperforms methods published over the past two years, while reducing both vision-encoder computation and the number of visual tokens processed by the LLM. We show that this result reflects a limitation of existing evaluations: most benchmark questions remain answerable from coarse visual information and rarely test localized fine-grained evidence. We introduce ResizeBench, a benchmark of 1,458 resolution-critical samples designed to measure this missing regime. Each sample is associated with a verified resolution barrier, the image-area budget at which its answer-relevant evidence becomes unrecoverable through resizing. We mine candidate barriers from a sweep of 387K samples and verify them using two frontier VLMs, answer filtering, and factuality checking. Across 15 token-compression methods, three VLM backbones, and 7.2M generated answers, the standard-suite and ResizeBench leaderboards exhibit weak agreement. Resizing performs strongly on standard benchmarks but degrades sharply when answers depend on localized detail, where selective compression provides substantial gains. The highest-ranked compression method also varies across backbones, indicating that token compression remains an open problem with substantial method–model dependence and no universally effective strategy. These results establish resizing as an essential matched-budget baseline and resolution-critical evaluation as a necessary complement to standard VQA benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.