Evaluating VLMs' Sensitivity to Image Resolution and Detail Level
Abstract
Visual Language Models (VLMs) have achieved remarkable success across diverse tasks, yet they struggle with high-resolution inputs where critical information resides in small regions or detailed, cluttered scenes. While several approaches address this limitation, a systematic understanding of why models fail at high resolutions is lacking. We introduce a controlled evaluation framework that disentangles resolution-related performance degradation from task difficulty through semantics-preserving transformations. We propose two simple metrics: Area Under the Scaling Curve (AUSC), which quantifies scaling robustness independent of baseline accuracy, and Prediction Variance Score (PVS), which measures resolution-induced prediction instability. Through comprehensive experiments across four model families and five benchmarks, we investigate architecture-dependent resolution sensitivity in relation to three pipeline mechanisms: (1) information loss from downsampling at vision token limits, (2) tokenization and positional sensitivity under changes in patch layout and aspect ratio, and (3) attention dilution as token counts increase. Our analysis reveals that widely used VLM architectures suffer from performance drops when processing high-resolution images, with degradation patterns varying systematically by architectural family. We provide actionable insights for model architecture design and data augmentation strategies to mitigate these limitations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.