Query Efficient Black-box Adversarial Attack via Dataset Distillation
Abstract
Score-based black-box attacks seek effective adversarial perturbations under limited access to the victim model, making query efficiency a central challenge. Existing methods often improve query efficiency by incorporating prior information, such as structured search priors or surrogate models to guide the search toward promising perturbation directions. However, the effectiveness of such guidance depends strongly on the quality and availability of the underlying prior. To provide a richer and alternative source of guidance, we propose DDAttack, which leverages dataset distillation (DD) as prior information for gradient-free adversarial search. DD condenses class-discriminative information from large datasets into a compact set of synthetic examples, providing informative input-space cues that can guide adversarial optimization. DDAttack exploits this distilled information through two complementary stages: distillation-guided initialization, which constructs a promising starting perturbation from distilled examples, and distillation-guided search, which aggregates promising candidate perturbations to guide subsequent updates. Extensive experiments on adversarially trained models demonstrate that distilled examples provide effective prior information for score-based black-box attacks. The source code is included in the supplementary material and will be publicly available upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.