When Marginal Stability Is Not Enough: Auditing Benchmark Conclusions under Composition Shift
Abstract
Multimodal benchmarks aggregate performance across heterogeneous evaluation contexts, yet it remains unclear how far their conclusions extend beyond the empirical mixture of those contexts. We introduce Composition-Conditional Reliability Auditing (CCRA), which keeps questions and model responses fixed and varies only the weights of observed context cells under bounded total-variation (TV) budgets. We audit four multimodal large language models (MLLMs) on MMAD, an industrial anomaly-detection benchmark, and a fixed three-benchmark panel spanning visual robustness (MVI-Bench), multidisciplinary reasoning (MMMU), and spatial reasoning (MMSI-Bench). Within the audited bounds, a one-percentage-point score change requires moving as little as 1.4–3.7% of evaluation mass on MMAD, MVI-Bench, and MMMU, whereas MMSI-Bench remains within 1.4 points. At the prespecified +1pp selection margin, joint reweighting reaches the endpoint in 6 of 12 native MMAD pair–prompt settings while matched product-only, task-only, and additive spaces do not; one of six MMMU pairs also reaches the endpoint. Larger challenger leads are substantially less common, and no five-point lead is reached within the audited scopes. Separately, on MMAD, joint stress increases held-out error among accepted answers by 6.9–9.7 percentage points for acceptance rules selected on development data and then frozen. These results show that benchmark robustness is conclusion- and margin-dependent, motivating reports that state the composition conditions supporting each score, model-selection, or acceptance conclusion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.