acceptodds
Under review as a conference paper at ICLR 2027

Machines See Better When They Imagine: Diffusion Imagination Reveals Better Visual Priors for Multimodal Reasoning

Abstract

Multimodal large language models (MLLMs) are typically evaluated with human-readable images, assuming that representations optimized for human perception are also optimal for machine reasoning. We challenge this assumption by studying whether intermediate images along a diffusion denoising trajectory, which acts like machine imagination, can serve as effective visual priors for visual question answering (VQA). Given an input image, we perform diffusion inversion and reconstruction under different conditional priors, and further derive residual difference maps that highlight question-induced visual changes. Evaluated across top tier MLLMs on our proposed benchmark DV-Bench, the integration of diffusion visual priors consistently improve MLLMs' performance, especially on reasoning-heavy tasks. Our analysis shows that the gains are not only driven by a universal denoising step or conditional priors, but also by task type, MLLM and diffusion backbone. The same diffusion prior can bring significant improvement for one MLLM while harming another, revealing a model-dependent “meat/poison” effect. These results suggest that diffusion priors as visual representation have great potential for enhance the visual reasoning of different MLLMs. The source code will be fully released upon paper acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.