WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Abstract
In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual diversity needed to handle open-ended visual inputs. We present WorldBench, a challenging and visually diverse reasoning benchmark to evaluate Multimodal Large Language Models (MLLMs). We build a taxonomy of thousands of visual concepts across multiple domains (e.g., living things). Guided by this taxonomy, we curate a broad collection of images from search engines and existing datasets to comprehensively represent the visual world. Through structured trial-and-error, we manually design challenging questions that frontier MLLMs fail to answer. Quantitative and human evaluations provide complementary evidence that WorldBench is among the most visually diverse benchmarks we evaluate, while each evaluation has distinct limitations. Evaluating 15 MLLMs on WorldBench reveals weaknesses in visual understanding: even the strongest model reaches only 64.0% accuracy, while some models perform marginally above chance level. We hope our work highlights the importance of visual diversity in building multimodal benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.