The Noise-Transmission Burden of Diffusion Network: Why Image Prediction Can Outperform Noise Prediction
Abstract
Clean-image prediction (-prediction) and noise prediction yield equivalent population-optimal denoisers, yet can perform very differently in finite diffusion networks. We relate this gap to a noise-transmission burden: forming a clean-image estimate from a noise predictor requires cancellation of corruption carried by the input. Under Gaussian corruption, a rank- linear encoder imposes a denoising-error floor for bare noise prediction even with unrestricted downstream computation, while image-prediction risk remains bounded. Analytic output preconditioning cancels this contribution and yields an image-prediction-equivalent learned branch; EDM suppresses it at high noise, with its learned target approaching a scaled clean image. Without input compression, and assuming exact affine recovery of the noisy patch, we characterize the minimum hidden width needed to approximate the population denoiser to a chosen accuracy. This width is determined by the patch dimension and the covariance spectrum of the denoiser's residual beyond its best affine approximation from that patch. Targeted JiT experiments on ImageNet examine these mechanisms. Learned output shortcuts substantially improve generation under the tested objective. Across two resolution/patch settings, widening improves generation quality and simultaneous affine access to noisy-input and clean-image information. Together, these results identify input compression and simultaneous input/denoiser accessibility as two concrete architectural mechanisms behind prediction-target differences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.