acceptodds
Under review as a conference paper at ICLR 2027

Evaluating the Gap Between VLMs and Specialized Vision Models at Visual Understanding

Abstract

Despite impressive improvements in general visual question-answering capabilities, Vision-Language Models (VLMs) continue to struggle at complex spatial reasoning tasks (e.g., comparing the relative positions of objects in a scene or identifying the direction of camera motion given two successive frames in a video) compared to specialized computer vision models (e.g., object detection/segmentation models). We perform a rigorous empirical evaluation to study this performance gap, focusing on two kinds of structural representations: (1) involving a single image or multiple images of unrelated scenes (e.g., bounding boxes), and (2) those involving multiple images of the same or similar scenes (e.g., correspondence points). The key challenge is that specialized models cannot directly solve realistic visual reasoning tasks. We propose to incorporate structural representations predicted by specialized models into the VLM's Chains-of-Thought (CoT), compared to incorporating the VLM's own predictions of the same structure. We find that using specialized models substantially outperforms using the VLM's own predictions at downstream reasoning tasks, rigorously establishing the limited spatial understanding of VLMs. Furthermore, we consider using supervised finetuning or reinforcement learning to improve the VLM's visual understanding capabilities; we find that in some cases these strategies can improve performance, but the gap compared to specialized models persists. Our experiments also provide insights into which kinds of structural representations are useful for improving VLM reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.