SPAR: Self-Referential Prompt-Conditioned Attention Reweighting for VLM Decoding
Abstract
Vision-language models often fail on fine-grained visual queries even when the required evidence is present in the image. We show that this failure arises from a decoder-side routing problem rather than a lack of visual perception. Across VLM families, early and final decoder layers fall into a Prompt-Invariant Sink Regime, where attention repeatedly collapses onto stable visual sink tokens across prompts. In contrast, middle layers form a Prompt-Conditioned Evidence Window, where attention is less sink-dominated and more responsive to the input query. Motivated by this depth asymmetry, we introduce SPAR, a self-referential and training-free attention reweighting method for frozen VLMs. SPAR estimates a region of interest from the model's own middle-layer visual attention and uses this prompt-conditioned evidence signal to guide visual attention during decoding. Across multiple benchmarks and VLM families, SPAR improves answer accuracy while increasing evidence retention and reducing mid-to-final evidence drop. Ablations over ROI construction, random controls, and layer choice show that the gains come from prompt-conditioned middle-layer evidence rather than generic attention perturbation or early-layer attention cues. Our results suggest that reliable fine-grained VLM decoding requires preserving evidence-bearing middle-layer signals through later decoder computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.