Think-Then-Verify: Decoupling Tool Verification from Visual Reasoning
Abstract
Multimodal large language models (MLLMs) can improve visual understanding and reasoning with external tools such as code interpreters. However, when tool use is tightly coupled with visual reasoning, a model must decide whether and how to use tools while pursuing a multi-step solution. On complex problems, these intertwined decisions can lead to unproductive exploration or code checks with answer-irrelevant hypotheses, reducing the effectiveness of verification and ultimately harming answer quality. We propose Think-Then-Verify (TTV), a phase-separated paradigm that decouples executable verification from answer formation. TTV first elicits tool-free reasoning and a draft answer, then extracts explicit, checkable hypotheses and invokes external tools solely to evaluate them. The resulting evidence determines whether the draft answer is retained or revised. By separating reasoning from verification, TTV makes each tool call accountable to a stated hypothesis rather than to an evolving interleaved trajectory. Across open- and closed-source MLLMs, TTV substantially increases targeted verification coverage and improves visual-reasoning accuracy over the interleaved tool-use baseline. We further construct a supervised training corpus for hypothesis extraction and code-based verification, and our fine-tuned checkpoints improve under TTV and transfer verification-oriented tool use to the interleaved tool-use baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.