acceptodds
Under review as a conference paper at ICLR 2027

PrimeFocus: Harnessing Intrinsic Regularities for Test-Time Policy Optimization in Vision-Language Models

Abstract

Test-time reinforcement learning (RL) adapts vision-language models (VLMs) on unlabeled test data by rewarding agreement with a majority-voted pseudo-label. However this method has two weaknesses. First, the reward depends only on the final textual answer, so a consensus that arises from language priors is reinforced as strongly as one supported by the image. Second, training runs for a fixed number of steps, although the best stopping point differs across test sets and no labeled validation set is available. We propose PrimeFocus, a test-time policy optimization method that adds intrinsic signals of the policy to this recipe. For the reward, consensus-anchored attention sharpening derives three annotation-free terms from the attention of answer tokens to image tokens: the share of attention on the image (visual enhancement), the concentration of that attention (visual focus), and the concentration of the attention map averaged over all responses that give the same answer (group visual consensus). For stopping, we track the cumulative deviation of the policy's mean token entropy from its exponential moving average, a statistic motivated by prequential estimates of epiplexity, and terminate training when it plateaus. On three visual recognition and four visual question answering benchmarks with InternVL3-2B and InternVL2.5-4B, PrimeFocus improves the mean accuracy of TTRV under the same backbone by 2.27and 2.42. Ablations and diagnostics examine each reward term, attention-level failure modes, and the behavior of the stopping rule.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.