CSRA: Decomposing Native Visual QK Alignment in Vision–Language Models
Abstract
Visual-language understanding in vision language models (VLMs) relies on effective cross-modal alignment. However, existing analyses based on attention maps primarily reveal where models attend, offering limited insight into how alignment forms and where failures originate. We show that cross-modal alignment is directly encoded in native query-key (QK) interactions and distinguish three sources of alignment failure: loss of textual semantic information during query projection, insufficient discriminability of visual target keys, and incorrect matching between queries and keys.Further analysis shows that even when queries preserve semantic structure and keys support target discrimination, their interactions can be dominated by a low-rank shared subspace associated with attention sinks, leading to deviations from semantic alignment. Based on this finding, we propose Common-Space Residual Attribution (CSRA), using singular value decomposition (SVD) to identify and remove the projection components responsible for this interference in QK interactions, thereby suppressing the influence of the sink subspace.Experiments on Qwen2.5-VL, LLaVA-1.5, and InternVL3 demonstrate improved target localization, with RefCOCO mIoU gains of up to 27.1 % over native attention. The estimated subspace also exhibits cross-dataset transferability. Our findings reveal how QK interactions shape cross-modal alignment and provide a unified framework for diagnosing potential alignment failures and guiding mechanism-driven alignment enhancement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.