acceptodds
Under review as a conference paper at ICLR 2027

Test-Time Training for Modality Order Consistency in Vision-Language Models

Abstract

We find that vision-language models are sensitive to a specific semantically irrel- evant change: the order in which the image and question are presented. Across three models on three core benchmarks, with additional evaluation on three further benchmarks, image-first prompting consistently outperforms question-first prompt- ing, revealing a repeatable modality order failure. We use this gap to design an order-consistent test-time training method. Our method substantially closes the modality-order gap across all evaluated settings. The stronger image-first branch is largely preserved, with observed improvements in several settings, hence boot- strapping both orderings toward mutual consistency. Activation patching localizes the ordering failure to a narrow mid-network region where representations diverge sharply between prompt orders. We find that the test-time training method repairs this misalignment across layers. Together, our results identify modality-order sen- sitivity as a circuit-level failure in VLMs and demonstrate that simple, asymmetric test-time adaptation can effectively mitigate it and even improve performance over the baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.