Same Causal Question, Different Lenses: Evidence-to-Answer Dependence in Vision-Language Models
Abstract
A vision–language model can answer a question correctly without using the image region that contains the answer. We introduce a region-anchored audit that connects visual evidence to answer-side computation and tests how much of that effect a proposed internal description reproduces. On compact-evidence questions, we compare answer-region masking with same-area control masking, restore clean hidden states during the masked run, and test candidate descriptions through paired clean-state restoration and masked-state writeback. Across Gemma3, Qwen2.5-VL, and LLaVA1.5, answer-region masks cause greater target-score loss than matched controls, and aligned hidden restoration selectively recovers that loss. At the tested interfaces, fitted sparse reconstructions, selected native channels, and example-shared low-rank subspaces do not reproduce their corresponding complete-reference effects. In contrast, a frozen LLaVA complete hidden-state interface restores 95.8% and removes 95.0% of the mean natural clean-to-mask score change on 60 examples with result-blind-reviewed masks, passing all four paired-bootstrap bounds. The audit thus identifies a shared visual answer dependence while distinguishing evidence-sensitive internal structure from a description that faithfully reproduces its measured answer effect.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.