acceptodds
Under review as a conference paper at ICLR 2027

CLOSING THE HALLUCINATION WINDOW IN DIFFUSION MODELS

Abstract

Diffusion models learn structure from noise, yet not always all of it, and this leads to hallucination. While the memorization and generalization regimes of their training dynamics have been well studied, we study how hallucination fits into these dynamics. We believe that by better understanding how hallucination occurs from an optimization perspective, we can better design regularization methods and, by extension, architectures for diffusion models that could offset hallucination during training. We find that denoising loss has an elbow, i.e., it plateaus significantly before the model actually generalizes . Consequently, hallucination can be distinguished from generalization only as an asynchrony in the optimization dynamics when the denoising loss is hierarchically resolved and rescaled. This gap, which we call the hierarchical lag, is the training dynamics signature of hallucination. We further show for a diffusion transformer trained on a random hierarchy model how a simple regularization to decrease this unintended “lag” can accelerate generalization and shrink the hallucination window in training time by an order of a magnitude.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.