Markov Proposals: Iterating Visual Attention for Vision–Language Alignment
Abstract
Despite rapid progress, Large Vision-Language Models (LVLMs) still generate descriptions unsupported by the visual input. One possible contributor is insufficient use of visual evidence during generation, involving both the amount of attention allocated to image tokens and its distribution across image regions. We propose a training-free method that guides the language decoder's attention using a spatial prior derived from the vision encoder. We interpret the encoder's self-attention as a weighted directed graph over image patches as nodes and propagate attention through this graph as a discrete-time Markov chain. The resulting Markov proposal incorporates indirect relationships among nodes to expose spatial structure beyond direct attention connections. We blend this image-dependent prior into the decoder's image-token attention at selected layers and normalize the full attention row. This intervention increases total image attention while guiding its spatial allocation through a combination of the shared proposal and the decoder's token-dependent attention. Our empirical analysis shows that a frozen LVLM can benefit from increased attention to annotated object regions. Controlled comparisons further show that Markov steering improves supported-token margins beyond proportional image-attention scaling, supporting a benefit from proposal-based redistribution of existing visual information. The method introduces no additional learned parameters and requires no auxiliary data or training. Experiments on LLaVA-1.5-7B and MiniGPT-4 demonstrate improvements across CHAIR, SPICE, BLEU, and POPE.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.