acceptodds
Under review as a conference paper at ICLR 2027

When More Thought Hurts: Calibrated EarlyStopping for Multimodal Reasoning

Abstract

Longer test-time reasoning is not monotonically beneficialfor multimodal large language models (MLLMs): under a shared-prefixcheckpoint protocol, an answer that is correct early can become wrongafter additional reasoning. We call these transitions harmful flips and usethem to study when extra inference has positive marginal value. AcrossVAPO-Thinker-3B and OpenVLThinker-7B, checkpoint oracles revealsubstantial per-example headroom, but this headroom does not implythat a learned router will transfer. We therefore evaluate a conservative,uncertainty-aware early-stopping family under locked target-benchmarkprotocols. For OpenVL, frozen routers remain non-inferior to 512-tokeninference while reducing mean reasoning tokens by 68.8% on MMBenchand 58.45% on SEED-Bench. The same approach yields little adaptivebenefit for VAPO, and on A-OKVQA a fixed 32-token budget is alreadyPareto-optimal, exposing an important calibration boundary. A separatevisual-intervention audit shows that original images remain causally use-ful relative to blank or counterfactual images, yet prompt-level re-viewdoes not reliably repair harmful trajectories. Our conclusion is deliber-ately narrow: checkpoint instability creates exploitable heterogeneity, butsafe savings require model- and distribution-calibrated stopping; neitheruniformly longer reasoning nor unconditional visual refresh is a reliablepolicy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.