PMF-Purify: Unifying Accuracy and Efficiency in Single-Step Pixel-Space Adversarial Purification
Abstract
Deep neural networks remain vulnerable to imperceptible adversarial perturbations. Adversarial purification provides a robust, plug-and-play defense by restoring perturbed inputs prior to classification without modifying the target model. While diffusion models offer powerful natural-image priors for this task, their iterative reverse process severely bottlenecks inference speed. Current acceleration methods largely attempt to compress these diffusion trajectories via distillation or guided few-step denoising. We argue for a paradigm shift: because adversarial examples inherently retain the vast majority of their original visual semantics, progressively regenerating them through a multi-step trajectory is fundamentally unnecessary. In this work, we introduce PMF-Purify, a single-step, latent-free adversarial defense operating in pixel space. Building on Pixel MeanFlow (pMF), PMF-Purify bypasses iterative sampling by formulating purification as a direct endpoint prediction. By utilizing a parameter-efficient LoRA adaptation—keeping the base generative prior and downstream classifier entirely frozen—it maps a perturbed intermediate state directly to the natural-image manifold in exactly one network evaluation (NFE = 1). Under the rigorous ImageNet-1K/ResNet-152 evaluation protocol, PMF-Purify achieves a state-of-the-art average robust accuracy of 78.6% across 11 diverse attacks. Strikingly, it accomplishes this with an inference latency of merely 3.72 ms per image, successfully unifying top-tier adversarial robustness with single-digit millisecond efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.