acceptodds
Under review as a conference paper at ICLR 2027

RoAR: Robust Test-Time Refinement with Unknown Priors for Vision-Language Models

Abstract

Test-time adaptation improves the zero-shot predictions of vision-language models using only the unlabeled test pool, and its gains are measured on pools whose priors are fixed: classes are roughly balanced, every class in the label set is present, and every image belongs to the label set. A deployed model receives a pool whose priors are unknown. We show that strong gains on the original benchmark pools need not persist when their class composition changes. The strongest evaluated baseline on the original pools falls below zero-shot in mean accuracy at the strongest tested level of class imbalance, class absence, and open-set contamination. We release an evaluation protocol that turns the class composition of a test pool into controlled prior axes, and we propose **RoAR** (Robust Anchored Refinement), a transductive test-time refinement for pools with unknown priors. It pairs a vision-language model with a self-supervised encoder, screens potentially unknown images before aggregation, and holds a selectively reweighted zero-shot anchor fixed during refinement. In two-view pooling, anchor confidence sets the total prototype exponent, while geometric scores determine the relative weights of the prototype views. With a single configuration shared across all datasets and prior settings, RoAR matches the strongest existing method on the original pools, improving over zero-shot by 7.7 percentage points on ImageNet and ten cross-domain datasets. Under severe class imbalance or absent classes, where that method drops below zero-shot, RoAR outperforms it by about 8 percentage points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.