acceptodds
Under review as a conference paper at ICLR 2027

What Happens Before Recognition? Background Effects on Target Structure in High-Resolution Vision-Language Models

Abstract

Fine-grained target recognition remains challenging for vision language models (VLMs) in high resolution images with extensive background. Existing work has mainly treated background interference as a target-search problem, while recent work points to an overlooked effect on changing the target itself. However, without separating and measuring these two background effects, the source of recognition failure remains unclear, obscuring the true bottleneck and limiting existing remedies to partial recovery. In this paper, we introduce a target-fixed causal framework that separates background interaction in the vision transformer (ViT) from background competition in the LLM and traces their consequences for recognition. Results show that changes in target representations induced by background contribute substantially more to recognition failure than token competition. We identify an attention fragility mechanism in early ViT blocks, where limited background interaction with target is sufficient to disrupt the target structure relevant for recognition. We further uncover LLM's compensatory localization response: as background grows, target selectivity increases despite declining recognition. Finally, we use layerwise repair and reveal two distinct propagation pathways, completing a causal link from upstream representation damage to downstream answer formation. These results establish target structure preservation as a missing dimension of VLM robustness, providing a mechanistic basis for both architecture design and evaluation under large visual contexts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.