acceptodds
Under review as a conference paper at ICLR 2027

Can Image Generation Benefit Visual Reasoning?

Abstract

If language models can "think in words", can multimodal models "think in images"? Recent advances in visual generation raise this possibility, but whether generated images can reliably support multimodal reasoning remains largely untested. We conduct a systematic study of visual generation for multimodal reasoning, spanning both modular VLM-Generator pipelines and unified models across diverse benchmarks. We find that, in common settings, incorporating visual generation does not improve over text-only reasoning, with gains confined to limited conditions. We find two main limitations. First, current models struggle to use generated images as effective intermediate reasoning steps: visual information that appears useful does not always translate into better downstream reasoning. Second, test-time scaling mainly benefits text reasoning, while scaling visual generation brings little improvement. Overall, current VLMs remain much stronger at reasoning in text than reasoning with generated images as intermediate evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.