acceptodds
Under review as a conference paper at ICLR 2027

Why Robust Classifiers Can Generate: Hidden Denoising structure

Abstract

Adversarially robust neural networks, while designed for classification, exhibit surprising generative capabilities when appropriately probed. We provide a theoretical framework explaining this phenomenon by connecting adversarial robustness to implicit denoising structure. Building on established results that robust training drives Jacobians toward low-rank solutions, we demonstrate that the Gram operator functions as an implicit denoiser, selectively preserving dominant directions while suppressing noise in orthogonal directions. We provide a mechanistic explanation for emergent generative capabilities in robust classifiers, validated through a diagnostic probing procedure that makes the denoising operator empirically observable. Our results establish a connection between discriminative robustness training and generative inference, showing that robust classifiers encode statistical priors that enable structured pattern generation without explicit generative objectives.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.