Good Rankings, Poor Choices? Rethinking Language-Only Vision–Language Model Selection
Abstract
Language-Only Vision–Language Model Selection (LOVM) aims to identify the Vision–Language Model (VLM) with the highest zero-shot accuracy on a target task using only textual information. We revisit the existing LOVM evaluation protocol and argue that ranking metrics, such as top-5 recall and Kendall's rank correlation, do not directly assess the performance of the selected model. We therefore adopt simple regret (SR), defined as the performance gap between the best and selected models, and introduce Outlier-Robust Normalized Regret (ORNR) to account for the varying performance ranges across tasks while reducing sensitivity to outliers. Building on this revised evaluation protocol, we propose Target-Conditioned Expected Source Regret (TC-ESR), a training-free method that selects the model with the lowest similarity-weighted mean regret across source tasks. Across 22 target tasks and 43 VLMs, TC-ESR achieves the lowest mean SR and ORNR among all compared methods. Separately, we identify two sources of target performance leakage in the prior ranking-based evaluation protocol and provide corrections that remove both. The code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.