VEC-TTA: Visual-Evidence-Calibrated Test-Time Adaptation for Compact Multimodal Language Models
Abstract
Test-time computation provides an opportunity to improve multimodal large language models (MLLMs) beyond their pretrained capabilities.However, reinforcing model predictions does not ensure that the resulting answers are driven by query-relevant visual evidence.Models can favor answers elicited without visual evidence even when the original image is present, and retain original answers after critical visual evidence is altered.To address these limitations, we propose Visual Evidence Calibrated Test-Time Adaptation (VEC-TTA), an instance-specific adaptation framework for compact MLLMs that performs temporary low-rank updates while keeping the backbone frozen.Specifically, VEC-TTA comprises three complementary modules:(1) Query-Conditioned Evidence Elicitation (QEE), which constructs pseudo-supervision from the current image–query instance;(2) Invariant Evidence Consolidation (IEC), which supplements answer fitting with visual consistency, jointly fitting pseudo-targets and aligning predictions across online content-preserving views; and(3) Counterfactual Evidence Calibration (CEC),which constrains persistent answer support by encouraging stronger support for the same pseudo-target when visual content is present than when it is removed.Experiments on two compact MLLMs across multiple benchmarks show consistent improvements over representative test-time methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.