acceptodds
Under review as a conference paper at ICLR 2027

Think-Then-Verify: Decoupling Tool Verification from Visual Reasoning

Abstract

Multimodal large language models (MLLMs) can improve visual understanding and reasoning with external tools such as code interpreters. However, when tool use is tightly coupled with visual reasoning, a model must decide whether and how to use tools while pursuing a multi-step solution. On complex problems, these intertwined decisions can lead to unproductive exploration or code checks with answer-irrelevant hypotheses, reducing the effectiveness of verification and ultimately harming answer quality. We propose Think-Then-Verify (TTV), a phase-separated paradigm that decouples executable verification from answer formation. TTV first elicits tool-free reasoning and a draft answer, then extracts explicit, checkable hypotheses and invokes external tools solely to evaluate them. The resulting evidence determines whether the draft answer is retained or revised. By separating reasoning from verification, TTV makes each tool call accountable to a stated hypothesis rather than to an evolving interleaved trajectory. Across open- and closed-source MLLMs, TTV substantially increases targeted verification coverage and improves visual-reasoning accuracy over the interleaved tool-use baseline. We further construct a supervised training corpus for hypothesis extraction and code-based verification, and our fine-tuned checkpoints improve under TTV and transfer verification-oriented tool use to the interleaved tool-use baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.