acceptodds
Under review as a conference paper at ICLR 2027

Beyond Online Accuracy: What Persists After Continual Test-Time Adaptation of Multimodal LLMs on Multiple-Choice VQA?

Abstract

Continual test-time adaptation (CTTA) lets multimodal large language models (MLLMs) update from unlabeled test streams, but it is usually evaluated only by online accuracy. This can confound the effect of adaptation with the inference procedure and does not reveal what the model retains after the stream. We introduce a two-stage benchmark for multiple-choice VQA that first measures online performance under paired presentation histories, then freezes each adapted model and evaluates it on shared held-out questions against a frozen model using the same inference procedure. Across ten published TTA and CTTA methods on Qwen2.5-VL, adapted models rarely outperform their matched frozen baselines, while temporary interface changes can leave persistent, order-dependent answer biases. Controlled updates further show that retained gains depend on the information available in the update target. Based on this finding, we propose Deliberate-then-Consolidate (DtC), which uses chain-of-thought (CoT) inference online and distills its CoT answer distribution into direct predictions. On Qwen2.5-VL-7B, DtC reaches the strongest frozen online inference and, on the benchmark split, improves clean held-out direct-answer accuracy by 2.1 points over frozen PA. Across four question splits, retained gains vary with the advantage of CoT over direct inference, a distinction that online accuracy alone does not reveal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.