Attention Dilution: Two Statistical Failure Modes of Visual Token Inflation
Abstract
Multimodal language models encode each image as hundreds to thousands of visual tokens. As this count grows, instruction following and fine-grained retrieval degrade, while prune and merge heuristics succeed or fail in ways that appear model dependent. A diffuse attention map need not imply a wrong answer, and no method explains when token count alone causes which failure. We develop a statistical theory of these phenomena. Softmax attention on a two population visual context fails in two ways. Evidence mass on the tokens that carry the answer decays algebraically past a budget set by the margin, and a typical query retains systematically less mass than the average one. Top 1 retrieval of a visual needle collapses at a much larger, exponentially separated scale. When the mass budget lies below the retrieval budget, a diluted yet correct window opens in which attention is already diffuse while the needle remains the argmax. The same split classifies queries by which budget binds and maps when score pruning or covering merges can work for a single frozen attention operation. Monte Carlo studies, a trained attention layer, a CLIP proxy, frozen Qwen3-VL decoders, and a LLaVA-1.5 encoder check instantiate the two laws. Generation accuracy on a referring scene index probe crosses one half in the predicted window once the query is read at the first generated token and the margin is locked from logits alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.