acceptodds
Under review as a conference paper at ICLR 2027

Blind Spatially-Varying Denoising in Image Diffusion Models

Abstract

Image diffusion models are usually trained with a single noise level for the whole image, so denoising can only proceed uniformly over all pixels at once. The few prior works that support spatially-varying denoising require an input noise map that specifies the noise level at every pixel location. In this paper, we propose blind spatially-varying denoising, which removes this requirement. We remove the timestep input of a pixel diffusion transformer and train it on spatially-varying noise. The model predicts the clean image, so subtracting its prediction from the input leaves the noise, and the size of this remainder gives the noise level of each patch. Any pixel-space model can measure these levels. Only one trained on spatially-varying noise can also denoise patch by patch. We show that we can add this ability to pretrained models by finetuning, including a text-to-image model. The result is still a generator with almost no loss in quality. Painting noise onto an image then becomes a single way to restore and edit it. The model removes strong spatially-varying noise, fills noise-covered regions, extends borders, and edits a photograph wherever noise is applied, without requiring a mask or noise map.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.