When Are Diffusion Recovery Clocks Identifiable? A Spectral Measurement Protocol for Teacher-Forced Denoising
Abstract
During teacher-forced diffusion denoising, low-frequency structure typically becomes recoverable before high-frequency detail, but this ordering could reflect image statistics, latent decoder effects, or a generic model bias. Standard latent-versus-pixel comparisons cannot isolate these explanations because they change representation, objective, architecture, and data scale together. We show that, for power-law natural spectra measured in log-spaced bands, an affine data-spectrum clock is nearly indistinguishable from a monotone band-index clock. Guided by a Gaussian posterior-mean null, we develop a content-sensitive recovery protocol with calibration-strength tests and spectrum-breaking positive controls. On SD3, the data-spectrum clock predicts held-out recovery better than a restricted diagonal pullback proxy but not reliably better than band index on natural images; PixelDiT replicates the coarse-to-fine ordering. Under the prespecified population-average log-radial measurement, engineered non-power-law spectra separate the candidate clocks, although the result is not basis-invariant. We also show that high-frequency power can be draw-specific texture rather than recovered target content. These results recast coarse-to-fine recovery as an identifiability problem: a pattern supports a mechanism only when competing clocks are forced to disagree.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.