UPD: Universal Perturbation Distillation from Individual Adversarial Attacks
Abstract
Deep neural networks remain highly susceptible to perturbation-based attacks, which seek small input modifications that induce model failure. These attacks manifest as either individual or universal adversarial perturbations (IAPs, UAPs), where the former are designed for specific inputs, whereas the latter are input-agnostic. While the simpler setting of IAPs has seen rapid methodological progress, UAP advancements remain comparatively limited, as adapting methods to the universal setting is often nontrivial. In this work, we propose *Universal Perturbation Distillation (UPD)*, a domain-decoupled formulation for learning universal adversarial perturbations from off-the-shelf IAP methods. By treating individual adversarial examples as representation-level supervision, UPD leverages IAP techniques for the universal setting. We instantiate UPD on both LLM jailbreak and image classification settings, achieving and often surpassing state-of-the-art performance, with substantial improvements on robust models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.