acceptodds
Under review as a conference paper at ICLR 2027

Harnessing Image Question Dependence for Better VLM Test-time Reinforcement Learning

Abstract

Test-time reinforcement learning can adapt vision–language models (VLMs) to unlabeled target data, but relies on reliable self-generated supervision. We analyze consensus-based VLM test-time reinforcement learning across diverse VQA datasets and model sizes and identify two limitations. First, much of its improvement comes from answer normalization rather than content correction. Second, many erroneous responses fail to jointly use the image and the question, allowing consensus rewards to reinforce grounding errors. We propose TTIQ, a test-time reinforcement learning framework that derives supervision from image–question dependence. For each sampled response, TTIQ measures how token likelihoods change under controlled image and question ablations, aggregates the resulting dependence into overall and sustained response-level scores for reward construction, and uses the token-level signals for credit assignment. The reward favors responses that are jointly supported by both inputs and sufficiently confident, while token-level credit emphasizes grounded parts of the response. This shifts supervision from output popularity toward grounded and reliable responses. Across eight VQA datasets and multiple model scales, TTIQ achieves the highest average performance at every evaluated scale, generalizes across VLM families, and transfers to unseen datasets after adaptation on a single source dataset.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.