Think Before You Paint: Recursive Latent Reasoning for Diffusion Models
Abstract
Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze. In a diffusion model an intermediate conclusion can only be carried forward by drawing it into the image. We argue that this is what impedes reasoning, because a wrong guess that has been partly drawn is hard to undo. We propose Painter-Thinker (PaTh), which performs additional reasoning steps before constructing each denoised image: a small recursive network (the Thinker) refines a latent state that is never rendered, and its result conditions the next step of a frozen denoising model (the Painter). The Thinker is trained with the standard reconstruction loss, with no solver or verifier. PaTh solves 92.5% of hard MNIST Sudoku puzzles (prior best 75%) and 71.2% of extreme ones (prior best 4.1%), with 10M parameters against 82M for a standard diffusion model. It also improves on mazes, Queens, and CLEVR scenes with specified spatial relations, and its advantage grows with problem size. Diagnostic experiments show that the diffusion model corrects only mistakes that are directly visible in the image, while PaTh also corrects mistakes that have to be inferred. Together, these results suggest that what diffusion models lack on rule-governed data is not capacity but a place to think, and that supplying one opens a path toward generating data under increasingly complex constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.