Evolutionary Adversarial Attacks Meet LLMs: Semantic-aware Adaptive Optimization Against Vision-Language Pre-trained Models
Abstract
Targeted attacks against black-box vision-language models remain challenging because transfer-based optimization must simultaneously achieve precise semantic steering, cross-model transferability, and bounded image perturbations without access to victim gradients or internal representations. Although evolutionary algorithms provide a natural gradient-free solution, their search is typically semantically blind and relies on fixed optimization strategies. We introduce LEA, an LLM-driven evolutionary adversarial framework that turns conventional evolutionary search into a semantic-aware adaptive optimization process. LEA first constructs original and target semantic anchors to explicitly characterize the desired semantic transition. More importantly, rather than directly generating adversarial perturbations, LEA employs the LLM as a high-level controller that interprets the evolving search state and adaptively coordinates semantic region priors, operator selection, and search parameters. A heterogeneous evolutionary operator pool then performs the low-level gradient-free perturbation search, enabling complementary global exploration and local refinement. Experiments on SMALLCAP, ViECap, and LLaVA-13B demonstrate consistent improvements over representative black-box and transfer-based baselines. On SMALLCAP, LEA improves Target ASR from 39.7% to 53.2% over the strongest targeted baseline while increasing the average CLIP score from 63.78 to 68.76. These results show that high-level LLM reasoning can effectively control evolutionary adversarial search and expose the vulnerability of vision-language systems to targeted semantic manipulation under transfer-based black-box settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.