When Exact Likelihood Is Anti-Robust: Diagnosing and Fixing Diffusion Classifiers under Domain Shift
Abstract
Generative classifiers are widely believed to be more robust than their discriminative counterparts. We show that for diffusion classifiers that score classes by exact likelihood the opposite holds under domain shift, and that the failure is structural: it follows from how the likelihood weights noise levels. Decomposing the class-conditional negative log-likelihood across noise levels via the information-theoretic diffusion identity, we find that the fraction of discriminative margin retained under shift rises steadily with noise level. Low-noise terms do not merely degrade out of distribution: classifying with them alone falls below chance, while high-noise terms retain roughly half of their in-domain margin. The exact likelihood integrates uniformly in log-SNR, which under the model's schedule places about three times the uniform-timestep weight on the lowest-noise bucket, precisely where evidence is most fragile. On the same model and the same per-noise-level denoising errors, the three scores order monotonically, exact < uniform-timestep < band-restricted, on all six Rotated MNIST folds and on CIFAR-10-C noise corruptions, where the exact likelihood is at chance and the uniform-timestep score used in practice reaches 36%. The diagnosis implies a one-line fix with no retraining: restrict the evidence integral to a noise band chosen without access to the test domain. On six-fold Rotated MNIST (DomainBed protocol), band restriction lifts a joint diffusion classifier from 57.9% to 92.7% mean OOD accuracy (31% to 87% on the hardest fold). In-domain validation alone selects a biased band; shifting it upward in noise by a fixed offset, as the diagnosis predicts, matches leave-one-domain-out selection while using only single-domain data. Transferred unchanged to CIFAR-10-C, the same rule recovers noise-family accuracy from 36% to 78% and lifts the repaired classifier above a matched ERM ResNet-18 on mean corruption accuracy (82% vs. 75%), though not above one trained with AugMix. Where restriction hurts, on fog, contrast, and the natural shift CIFAR-10.1, the damage sits in the high-noise band, so the per-bucket decomposition predicts which shifts the fix helps. For the shifts we study, these results reverse the robustness narrative for exact-likelihood classification and reduce the repair to a one-line change in the evidence integral.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.