TPA: Early Text-Embedding Perturbations for Natural Adversarial Diffusion
Abstract
Natural adversarial examples show that classifiers can fail at test time even on images that are clearly recognizable to humans, not just on images modified with small input perturbations. A key challenge is to make adversarial examples work across different models without reducing image quality. Text conditioning influences which visual variations of a class the diffusion model explores. Keeping the text condition fixed may restrict adversarial search to only a subset of these variations. We propose Text-embedding Perturbation Attack (TPA), which temporarily perturbs text embeddings during early denoising and progressively restores the original conditions while guiding sampling with classifier gradients. This exploration changes both the denoising trajectory and the intermediate images on which adversarial guidance acts. On ImageNet-1K, TPA achieves higher black-box attack success rates than existing natural adversarial generation methods across all evaluated surrogate–victim pairs, while outperforming NatADiff on most learned image-quality metrics. It also achieves lower FID to ImageNet-A than both baselines. Matched ablations show consistent transfer gains from early conditioning perturbations across the evaluated black-box classifiers, including adversarially trained models. These results establish early text-conditioning exploration as an effective strategy for generating transferable natural adversarial examples and probing test-time vulnerabilities across classifiers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.