acceptodds
Under review as a conference paper at ICLR 2027

CAN LARGE MULTIMODAL MODELS THINK JUST ENOUGH? EVALUATING OUTPUT-SIDE ADAPTIVE REASONING IN LARGE MULTIMODAL MODELS

Abstract

Explicit reasoning can improve the performance of large multimodal models (LMMs) on complex visual understanding and reasoning tasks, but lengthy and unnecessary reasoning also incurs additional generation cost. Although recent output-side adaptive reasoning methods have shown potential for reducing unnecessary generation cost, their evaluation remains fragmented and inconsistent. In this work, we introduce MARE-Bench, a benchmark for systematically evaluating output-side adaptive reasoning methods for LMMs. MARE-Bench provides unified evaluation protocols across seven representative datasets spanning five evaluation dimensions: General Knowledge, Mathematical Reasoning, Logical Reasoning, Medical QA, and OCR. It covers seven representative adaptive reasoning methods and three LMM families. We analyze adaptive reasoning methods in terms of task performance, average output tokens, and reasoning behaviors, characterizing the performance–generation cost trade-off. Our experiments reveal several findings: (1) some methods outperform their base models, while others perform poorly; (2) FAST-GRPO achieves consistently high overall performance across the evaluated settings; (3) the effects of adaptive reasoning vary across evaluation dimensions, with larger performance gains on Mathematical Reasoning and Logical Reasoning tasks, smaller gains in General Knowledge and Medical QA, and performance degradation on OCR; and (4) the performance–generation cost trade-off of adaptive reasoning methods can vary across base models. We believe that MARE-Bench will serve as a reliable foundation for future research on efficient adaptive reasoning in LMMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.