acceptodds
Under review as a conference paper at ICLR 2027

Which Attention Heads in VLMs Should We Adapt at Test Time?

Abstract

Training vision-language models (VLMs) at test time has emerged as a promising approach to improving visual understanding and mitigating prediction failures during inference. However, existing methods largely predefine the model components to be adapted, risking inefficient parameter updates and unnecessary model drift, which are particularly problematic in the data-scarce test-time setting. We introduce TTT-Head (Test-Time Training of Attention Heads), a fine-grained TTT framework that automatically identifies tunable attention heads and performs instance-specific adaptation in VLMs. Experiments on seven benchmarks covering diverse VQA scenarios—including visual reasoning and compositional question answering, hallucination recognition, multimodal perception and reasoning, and optical character recognition—show that TTT-Head consistently improves the corresponding frozen baselines across all evaluated VLMs, with an average gain of 2.65 points on the LLaVA-1.5-7B backbone. We further demonstrate its generalizability across three widely used VLM families spanning 7B and 13B model scales. In-depth analysis reveals that suppressive interventions on identified key heads are more frequently associated with successful corrections than amplifying interventions, while tunable heads exhibit structured layer-wise preferences across different answer formats.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.