acceptodds
Under review as a conference paper at ICLR 2027

Measuring, Predicting, and Diagnosing Compression-Induced Failures in Vision-Language Models

Abstract

Visual-token compression reduces the inference cost of vision-language models (VLMs). It is usually evaluated by average accuracy, which does not show which correct answers become incorrect. We characterize these compression-induced failures (CIFs) through paired reference and compressed runs, defining a CIF as a sample answered correctly by the reference but incorrectly after compression. We focus on post-prefill KV-cache eviction across ten benchmarks, three model families, four controlled retention policies, and adaptations of six published strategies. Methods with nearly equal accuracy often break different requests. On DocVQA, two adapted attention-based eviction methods differ by 0.19 points, yet 17.5% of the dataset fails under exactly one of them. The failures concentrate on benchmarks whose answers rely on small visual details such as text and layout. The paired runs also make these failures predictable and repairable. Output-confidence features rank CIFs at AUROC up to 0.89 and support selective fallback in paired simulation. Restoring pruned tokens repairs most failures, and at tight budgets score-guided selection repairs more than random or spatial selection. These results support reporting paired failure counts alongside average accuracy when evaluating compressed VLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.