acceptodds
Under review as a conference paper at ICLR 2027

Perturbed Visual Keys and Values as a Contrastive Reference for Multi-Image Reasoning

Abstract

Multi-image reasoning requires multimodal large language models (MLLMs) to integrate visual evidence across images. However, how cross-image interaction shapes the visual representations used for generation remains unclear. We examine this interaction through visual Keys and Values, which govern how text queries attend to and aggregate visual information. We observe that cross-image interaction changes both components: disabling its induced change in either one lowers accuracy. Describing later-image Keys and Values as coefficients along directions defined by earlier images, we find that cross-image interaction makes more of these coefficients negative. Removing or reversing the negative ones progressively weakens support for the correct answer. This degradation serves as an effective reference branch for Key-Perturbed Contrastive Decoding (K-PCD) and Value-Perturbed Contrastive Decoding (V-PCD), two training-free methods that contrast original predictions against internally perturbed visual branches at inference time. Across seven models from three families and four multi-image benchmarks, K-PCD and V-PCD achieve the two highest average accuracies and outperform state-of-the-art baselines in cross-model average accuracy on three of four benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.