acceptodds
Under review as a conference paper at ICLR 2027

Not Blind but Silenced: Preserving Visual Evidence through Counter-Context Validation

Abstract

Multimodal large language models (MLLMs) can produce incorrect answers even when intermediate representations robustly favor the factual hypothesis. We characterize this failure through layer-wise evidence competition and identify strict silencing, in which established intermediate factual dominance is overturned before the final prediction. To preserve the influence of useful visual evidence, we propose Adaptive Counter-context Evidence (ACE), a training-free decoding framework that follows Challenge, Validate, Locate, and Preserve. ACE constructs a controlled counter-context to estimate contribution reliability, locates an evidence anchor, and selectively amplifies original-image contributions throughout the decoder suffix. The counter-context is discarded before answer generation, while reliability gates remain fixed and current contributions evolve with each decoding step. Across four MLLM backbones, ACE improves POPE-Adversarial accuracy by 6.37–8.97 percentage points and reduces CHAIR by 3.9–5.0 points, while maintaining MM-Vet performance within 0.1 point of greedy decoding. Functional diagnostics demonstrate stronger evidence identification, while ACE achieves four times as many strict-silencing rescues as uniform-suffix preservation on the fixed Qwen3 evaluation cohort.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.