acceptodds
Under review as a conference paper at ICLR 2027

EvidenceFlow: Inference-Time Suppression of Confounder Evidence in Vision-Language Models

Abstract

A vision-language model often affirms an object that is absent from an image when a related object is visible. Such errors are usually attributed to weak visual grounding, yet the visible content behind them is accurate and genuinely used: the failure is one of attribution rather than amount, so supplying more or less visual influence does not address it. We make this failure measurable per example. A visible cue is confounder evidence for an object query when removing it lowers the model's Yes margin more than every area-matched control removal, and the error it sustains is a context-supported false affirmation. In every evaluated model, removing co-occurring cues fixed in advance lowers the Yes margin more than the controls do on average, and restoration and blocking show that this support is concentrated in a narrow layer window at the prediction position. EvidenceFlow acts in that window at inference time: when the model answers Yes, it subtracts the positive component of the prediction state along a calibrated confounder direction, removing the cue's support for the queried claim while keeping the visual content. The direction and window come from 12-17 image pairs per model, one strength serves all models, and deployment uses only the original image and question with frozen weights. On POPEv2, EvidenceFlow improves InternVL3-8B accuracy by 8.7 percentage points, correcting 35.4% of false affirmatives while retaining 98.4% of correct ones, and accuracy improves in all nine model-benchmark cells. On Causal-HalBench with InternVL3-8B, replacing the calibrated direction with a random one at the same window and strength lowers correction from 46.2% to 1.1%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.