MMR-Bench V2: From Visual Information to Routing Decisions
Abstract
Multimodal large language models vary in capability and inference cost, motivating query-dependent routing. When do model diversity and visual information translate into better model choices? MMR-Bench V2 expands MMR-Bench to 44 models and 18 benchmarks, recording outcomes, execution status, and reference costs for every query–model pair. We examine model-pool composition, query information, and target-task adaptation. At the pool level, expansion increases answer coverage faster than routing quality. Compact global pools retain the quality of unrestricted task-conditioned policies within a predefined tolerance at medium and high validation budgets on the evaluated suite. At the information level, matched frozen-feature comparisons show that images can improve outcome prediction while worsening selection. Most prediction gains concern the mean success rate across candidates for a given query. Visual effects depend on task composition, and preferences for models that generally perform well on each task provide strong controls when those tasks are represented in training. For known-task adaptation with matched prior weights, concentrating target labels on candidates screened using other tasks can outperform tested broader uniform and sequential allocations at the same label budget. These findings favor learning task-level preferences and evaluating larger pools and richer inputs by their additional selection gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.