acceptodds
Under review as a conference paper at ICLR 2027

Enabling Visual Reasoning over Generated Images in Diffusion Models

Abstract

Reasoning has become central to LLMs, whereas diffusion models have largely been treated as visual generators. As image generation increasingly requires understanding complex and compositional prompts, existing diffusion models often rely on external LLMs or VLMs for reasoning, leaving instruction understanding decoupled from visual generation. We introduce Seenk, a visual reasoning paradigm for correcting mismatches between generated images and their text prompts. Given an image-description pair, Seenk determines whether any mismatches exist and, if so, generates a mask-annotated image to localize them. This image serves as an explicit visual thought, guiding an SFT-adapted editor to correct the identified mismatches. To equip diffusion models with this capability, we introduce two core visual reasoning modules: VR-Base and VR-Refiner. VR-Base is trained via SFT on VRB-Dataset, a visual reasoning dataset constructed through our data pipeline, to learn mismatch detection and visualization. VR-Refiner is trained via DPO on VRR-Dataset, a preference dataset constructed with a VLM, to further improve visual reasoning and mask localization accuracy. Experiments on visual reasoning and correction demonstrate that Seenk achieves competitive performance against strong open- and closed-source models, as well as pipelines that employ an LLM as an external reasoner. Further experiments show that, like LLMs, Seenk benefits from larger reasoning budgets, generating more visual thoughts to better handle complex prompts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.