acceptodds
Under review as a conference paper at ICLR 2027

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

Abstract

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual diversity needed to handle open-ended visual inputs. We present WorldBench, a challenging and visually diverse reasoning benchmark to evaluate Multimodal Large Language Models (MLLMs). We build a taxonomy of thousands of visual concepts across multiple domains (e.g., living things). Guided by this taxonomy, we curate a broad collection of images from search engines and existing datasets to comprehensively represent the visual world. Through structured trial-and-error, we manually design challenging questions that frontier MLLMs fail to answer. Quantitative and human evaluations provide complementary evidence that WorldBench is among the most visually diverse benchmarks we evaluate, while each evaluation has distinct limitations. Evaluating 15 MLLMs on WorldBench reveals weaknesses in visual understanding: even the strongest model reaches only 64.0% accuracy, while some models perform marginally above chance level. We hope our work highlights the importance of visual diversity in building multimodal benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.