acceptodds
Under review as a conference paper at ICLR 2027

Adversarial Scenario Attack: Query-Efficient Black-Box Generation of Natural Adversarial Examples

Abstract

Generating natural adversarial examples (NAEs) from a given image remains challenging under tight black-box query budgets, as many existing generative attacks rely on surrogate-model guidance or attack-specific training. We propose Adversarial Scenario Attack (ASA), a query-based black-box attack on image classifiers that searches over human-readable natural-language editing scenarios rather than pixel-level perturbations. ASA uses a multimodal large language model (MLLM) to propose image-dependent changes to backgrounds, weather, and surface appearance, and a generative editor to apply them to the original image. Winner–loser feedback guides subsequent proposals, while a Greedy Explorer composes complementary scenarios and retains only combinations that improve the attack score. Under a budget of 100 victim-model queries per image, ASA achieves the highest attack success rate on all ten evaluated ImageNet classifiers, including two vision–language models, without surrogate models or attack-specific training. On average, its success rate within 20 queries already exceeds that of every baseline at 100 queries, while it maintains competitive image-level transferability. Because each attack is expressed as a readable scenario, the discovered edits can also be reused: class-neutral scenarios that recur during search, such as light mist or dusk lighting, raise the average error rate to 2.0–3.3 times that of a control prompt when applied unchanged to images from all 1,000 ImageNet classes across 13 classifiers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.