Same Photo, More Evidence? Multi-Image VLMs Overcount Repeated Images
Abstract
Vision–language models increasingly answer questions about a set of images rather than one: the photographs on a product listing, the images attached to an incident report, the results of a web search. In such sets the same photograph routinely re-enters re-compressed, cropped or screenshotted. Existing multi-image benchmarks vary how many images a model sees and what they depict, but never whether an added image is a copy of one already present. The standard remedy, de-duplication by embedding similarity, does not settle this either: a re-upload of a photograph and a genuinely new photograph of the same object look alike to it, so any threshold either keeps the copies or throws away real evidence. We show that current multi-image VLMs count images rather than new observations. On a controlled spectrum that fixes image and token count, renderings of one photograph raise a model’s belief almost as much as genuinely new viewpoints; on real product listings, three re-uploads of one mis-attributed photo change up to a third of decisions, on open 3B–72B models and on API-served ones alike. Copies of a correct photograph inflate confidence as well, and stating in the prompt that copies count once does not help. Attention analysis places the effect not in how much the answer attends to the repeated images, which is flat across conditions, but in reinforcement among those images’ own tokens in the middle layers. The remedy is to remove the copies before the model sees the set. Provenance-Aware Evidence Aggregation (PAEN) scores each pair of images for provenance, asking whether they are renderings of one observation, by geometric verification: SIFT correspondences under a RANSAC homography, backed by a perceptual hash. This separates re-uploads from nearby viewpoints where semantic similarity cannot, even with a descriptor trained for copy detection. PAEN is training-free, returns four open architectures to within two points of their base accuracy with near-zero decision change while leaving genuinely new photographs untouched, and, applied before the images are sent, does the same for three models served only through an API (DeepSeek, GPT-5.6, Kimi). The signal multi-image aggregation needs is provenance, not similarity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.