acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Current Answer: Visual Content Transfer into VLM Text States

Abstract

When a vision–language model tolerates removal of visual states, has it retained image content or only computed the current answer? We separate three questions: which states control the answer, whether text states contain information beyond that answer, and whether visual states remain necessary. Our central test pairs real images with identical questions and annotated answers but different unasked object categories. Patching donor visual activations early versus late changes which image's unasked content is decodable from non-final text states. Across seven tested checkpoints and 500 disjoint pairs per checkpoint, the paired interactions are positive, including a model with no sustained visual–text deletion gap. Complementary opposite-answer patches reveal a depth-dependent shift of causal control from visual to non-final text states. Symmetric zero-out on 1,000 ScienceQA questions yields a positive first-recovery gap in six of seven checkpoints, but sustained ordering is architecture-dependent; open-ended TextVQA reverses the first-recovery ordering in two models. Position-controlled probes expose a substantial causal-order confound. These results support transfer of image-derived semantic content beyond the current answer into text residuals. They do not establish lossless image storage or unrestricted re-queryability: a visual-KV-blocked second-question test largely fails. Content transfer and visual-token dispensability are distinct properties.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.