acceptodds
Under review as a conference paper at ICLR 2027

Same Image, Different Ranking: Task Formulation Matters in MLLM Evaluation

Abstract

Multimodal Large Language Models (MLLMs) are typically evaluated using benchmark scores that are often interpreted as stable measures of model capability. In this work we demonstrate inconsistencies in the VLM/MLLM performance measures in popular benchmarks. Instead, rankings tend to align with benchmark identity. That is, tasks within a single benchmark, even those targeting different capabilities, yield similar model rankings. % Crucially, we identify task formulation, the textual definition of the task input and output format, as a source of such discrepancies. To study this, we construct controlled variants of benchmarks, holding the visual content and semantic goal fixed for each item, while varying the task formulation (binary, multiple-choice, or open-ended). We demonstrate that this change can substantially alter both model success on the same example and the resulting overall model rankings. A natural response is to evaluate tasks under multiple formulations, but this substantially increases evaluation costs. To mitigate this, we propose a \em formulation-aware subset selection based on Item Response Theory (IRT). This approach efficiently approximates multi-formulation evaluation using compact subsets, which we confirm recover the performance trends as the full matrix. We conclude that reliable MLLM evaluation must account for variation across task formulations, and offer an effective and efficient protocol to do so.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.