acceptodds
Under review as a conference paper at ICLR 2027

VEC-TTA: Visual-Evidence-Calibrated Test-Time Adaptation for Compact Multimodal Language Models

Abstract

Test-time computation provides an opportunity to improve multimodal large language models (MLLMs) beyond their pretrained capabilities.However, reinforcing model predictions does not ensure that the resulting answers are driven by query-relevant visual evidence.Models can favor answers elicited without visual evidence even when the original image is present, and retain original answers after critical visual evidence is altered.To address these limitations, we propose Visual Evidence Calibrated Test-Time Adaptation (VEC-TTA), an instance-specific adaptation framework for compact MLLMs that performs temporary low-rank updates while keeping the backbone frozen.Specifically, VEC-TTA comprises three complementary modules:(1) Query-Conditioned Evidence Elicitation (QEE), which constructs pseudo-supervision from the current image–query instance;(2) Invariant Evidence Consolidation (IEC), which supplements answer fitting with visual consistency, jointly fitting pseudo-targets and aligning predictions across online content-preserving views; and(3) Counterfactual Evidence Calibration (CEC),which constrains persistent answer support by encouraging stronger support for the same pseudo-target when visual content is present than when it is removed.Experiments on two compact MLLMs across multiple benchmarks show consistent improvements over representative test-time methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.