acceptodds
Under review as a conference paper at ICLR 2027

SPAR: Self-Referential Prompt-Conditioned Attention Reweighting for VLM Decoding

Abstract

Vision-language models often fail on fine-grained visual queries even when the required evidence is present in the image. We show that this failure arises from a decoder-side routing problem rather than a lack of visual perception. Across VLM families, early and final decoder layers fall into a Prompt-Invariant Sink Regime, where attention repeatedly collapses onto stable visual sink tokens across prompts. In contrast, middle layers form a Prompt-Conditioned Evidence Window, where attention is less sink-dominated and more responsive to the input query. Motivated by this depth asymmetry, we introduce SPAR, a self-referential and training-free attention reweighting method for frozen VLMs. SPAR estimates a region of interest from the model's own middle-layer visual attention and uses this prompt-conditioned evidence signal to guide visual attention during decoding. Across multiple benchmarks and VLM families, SPAR improves answer accuracy while increasing evidence retention and reducing mid-to-final evidence drop. Ablations over ROI construction, random controls, and layer choice show that the gains come from prompt-conditioned middle-layer evidence rather than generic attention perturbation or early-layer attention cues. Our results suggest that reliable fine-grained VLM decoding requires preserving evidence-bearing middle-layer signals through later decoder computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.