Perturbed Visual Keys and Values as a Contrastive Reference for Multi-Image Reasoning
Abstract
Multi-image reasoning requires multimodal large language models (MLLMs) to integrate visual evidence across images. However, how cross-image interaction shapes the visual representations used for generation remains unclear. We examine this interaction through visual Keys and Values, which govern how text queries attend to and aggregate visual information. We observe that cross-image interaction changes both components: disabling its induced change in either one lowers accuracy. Describing later-image Keys and Values as coefficients along directions defined by earlier images, we find that cross-image interaction makes more of these coefficients negative. Removing or reversing the negative ones progressively weakens support for the correct answer. This degradation serves as an effective reference branch for Key-Perturbed Contrastive Decoding (K-PCD) and Value-Perturbed Contrastive Decoding (V-PCD), two training-free methods that contrast original predictions against internally perturbed visual branches at inference time. Across seven models from three families and four multi-image benchmarks, K-PCD and V-PCD achieve the two highest average accuracies and outperform state-of-the-art baselines in cross-model average accuracy on three of four benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.