Image Keys and Values are Asymmetric in MLLM Text-to-Image Attention
Abstract
Multimodal Large Language Models (MLLMs) integrate visual information via text-to-image (T2I) attention in the decoder. While recent studies identify various T2I attention phenomena, the underlying signals driving this attention remain underexplored. We investigate this through a structural analysis of image Key and Value representations, decomposing each into a visual prior and an instance-aware residual. This uncovers two Key–Value asymmetries: within each layer, image Keys are prior-dominated while Values are residual-dominated; across layers, the residual grows in Values but fades in deep-layer Keys. A cross-distribution analysis further shows the Key prior is input-agnostic, retaining high similarity under extreme distributional shifts. This reveals a routing–content gap: deep layers carry rich instance-aware content in Values, yet route attention through Keys dominated by an input-agnostic prior. Notably, within the same model family, the larger model exhibits a greater instance-aware component in its image Keys, providing complementary support for residual amplification. We therefore propose Image Key Steering, a training-free approach that amplifies the instance-aware component of deep-layer image Keys. Across four MLLMs and ten benchmarks, Image Key Steering consistently improves base-model performance, with the largest gains on tasks requiring precise visual grounding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.