acceptodds
Under review as a conference paper at ICLR 2027

Rarity-Aware Vision-Language Benchmarking and Prompting

Abstract

Prompt learning for vision-language models is predominantly evaluated under balanced supervision, leaving the effects of class imbalance on recognition and transfer insufficiently understood. We introduce TailPromptBench, a benchmark for long-tailed prompt generalization that varies class imbalance under a fixed annotation budget and evaluates recognition across frequency groups, base-to-new transfer, and robustness to visual domain shift. This design separates unequal supervision from changes in training-set size. We further propose RAP, a rarity-aware training parameterization for prompt learning based on a shared residual context. Its contribution increases with class rarity, providing stronger auxiliary corrections for underrepresented classes without introducing class-specific parameters. Joint optimization allows rarity-dependent contexts to shape the prompt parameters retained for deployment. The residual is discarded at inference, preserving the original structure and cost without requiring frequency estimates for unseen categories. Experiments show stronger tail-class recognition and average base-to-new harmonic-mean gains of up to 2.8 percentage points, with benefits extending to shifted domains. Ablations support the importance of aligning residual adaptation with the actual class-frequency structure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.