acceptodds
Under review as a conference paper at ICLR 2027

When More Seeing Does Not Help: Understanding and Mitigating Hallucinations in Think-with-Images Models

Abstract

Think-with-Images has emerged as a promising paradigm for multimodal reasoning. By acquiring visual evidence through tool calls (e.g., cropping and zooming), these models are widely expected to ground their reasoning in the image and mitigate visual hallucinations. In this paper, we revisit this paradigm and find that more tool invocations fail to consistently reduce hallucinations. Our analyses uncover a pronounced visual recency bias: the model over-relies on the recent visual observation while neglecting salient evidence from preceding tool calls, leading to hallucinations. Motivated by these findings, we propose Rebalancing Attention over Visual Evidence (RAVE), a lightweight, training-free strategy that recycles excessive attention from visual tokens in the latest tool output and reallocates it to historical visual contexts. Extensive experiments across six models and fifteen benchmarks demonstrate that RAVE not only mitigates hallucinations but also enhances multimodal perception and reasoning capabilities. These results show that reliable multimodal reasoning hinges not on invoking more tools, but on how to effectively integrate them along the reasoning chain. The code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.