Ranking Black-Box LMMs Without Labels via Visually Grounded Cross-Model Agreement
Abstract
Ranking large multimodal models (LMMs) on a new task typically requires labeled evaluation data, yet in practice target data may be unlabeled and candidate models may be accessible only as black boxes. Existing unsupervised ranking methods commonly rely on model internals, labeled proxy data, or repeated generations, which are costly and can favor models that are consistently wrong. We ask whether candidate LMMs can instead provide supervision for ranking one another. We introduce Cross-Model Agreement with Visual Evidence for Ranking (MAVER), a label-free and training-free framework for LMM ranking. Given one response from each candidate model, MAVER constructs an NLI-weighted response graph and characterizes agreement using semantic-cluster prevalence and within-cluster connectivity. However, agreement alone can be misleading when multiple models share the same language prior and converge on the same incorrect answer. MAVER therefore intervenes on the image while keeping the question fixed, and measures the resulting semantic response change to quantify how strongly each response depends on the visual input. This visual-dependence signal is used to refine the support assigned to each semantic answer, while cross-model semantic entropy gives greater weight to image–question pairs that better distinguish the candidate models. Across 32 LMMs and 15 datasets, MAVER outperforms 14 comparison methods in average ranking performance, showing that cross-model agreement becomes a reliable signal for label-free LMM ranking when complemented by evidence of visual dependence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.