Hard to Unsee: Source Addressability After Multimodal Integration
Abstract
A system that integrates several sources often learns which one to trust only afterwards, when the original evidence can no longer be re-read. We ask what a vision-language model can still do with a late, source-conditioned instruction once only an intermediate representation remains. An image and one or two texts disagree about one attribute; an inserted summary records or omits which source said what; access to the evidence is kept or removed by a verified attention mask; only then is a source queried by name or declared unreliable. After a source-neutral summary, Qwen3-VL-8B and Pixtral-12B do not retrieve what a named non-default source reported above a shortcut ceiling, and a late reliability statement is followed only when it agrees with a persistent, image-favoring default. Carrying arbitrary source handles through integration, without binding any value to them, restores most named-source retrieval; handles introduced after the bottleneck do not; answers follow a handle when its source assignment is permuted; the effect survives eviction of the evidence from the cache; and blocking attention to the carried handle at question time removes it while a matched sham does not. The hierarchy replicates in Pixtral-12B and on controlled natural images, is weaker for visual than for text targets, and reliability-conditioned selection lags retrieval only in Qwen3-VL-8B. In these settings, later source-conditioned access depends strongly on a referent carried through integration, while explicit source-value binding provides additional support for reliability-conditioned selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.