Readability Is Not Reliability: Semantic Readouts in Unconditional Diffusion Models
Abstract
Semantic probes are widely used to interpret diffusion representations, yet accurate decoding alone does not guarantee that the underlying predictive relationship persists across noise realizations and generative states. In this work, we systematically investigate semantic readout reliability in unconditional diffusion models by distinguishing mere readability from reproducibility under noise resampling and predictive reuse along reverse trajectories. Through theoretical analysis, we characterize why readable representations can yield unreliable readouts: changes in the optimal predictive mapping undermine cross-state reuse, while nuisance variation and insufficient identifiability limit reproducible readout recovery. For linear readouts, we derive exact risk decompositions to quantify decoder mismatch and the nuisance contribution from averaging independent noise realizations, establishing sufficient conditions for predictive-subspace recovery and local reuse. Empirical evaluations on CelebA-HQ latent diffusion and CIFAR-10 pixel-space diffusion reveal readable representations possessing with unstable readout geometries. Furthermore, an independent analysis of 5,000 reverse trajectories demonstrates that accurate attribute decoding frequently coexists with severe transfer failures of an unchanged readout. Our results expose the critical gap between semantic accessibility and readout persistence, proving that reliable diffusion interpretability demands evidence beyond static probe accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.