EvoNAT: Natural Language-Guided Evolution of Adversarial Images for Large Vision-Language Models
Abstract
Natural adversarial images induce semantic errors in large vision-language models (LVLMs) while remaining visually natural. Existing source-image attacks typically search for local pixel perturbations around benign images, tying the search to the source and limiting the visual content that can be explored. We introduce EvoNAT, an output-only black-box framework that evolves structured visual descriptions in semantic space using image-captioning feedback. EvoNAT generates complete adversarial scenes without requiring a source image. We further extend this approach to source-conditioned editing with EvoNAT-Edit, which evolves source-grounded edit specifications and renders each candidate from the original image. Both branches use image-captioning feedback to optimize semantic deviation, while visual question answering (VQA) is held out to evaluate cross-task transfer. Across seven open-weight and proprietary LVLMs, EvoNAT achieves relative gains of 53.4% and 52.8% in the seven-model mean cosine-based IC-Attack and VQA-Attack scores, respectively, over the strongest baseline on each task. Its normalized text-only LLM Judge scores also show relative gains of 2.0% on image captioning and 4.2% on VQA. The generated images achieve a 1.2% higher MUSIQ score and a 14.3% higher human naturalness rating than the strongest source-free baselines on the respective metrics. In source-conditioned editing, EvoNAT-Edit achieves relative gains of 36.2% and 22.7% in the mean cosine-based IC-Attack and VQA-Attack scores, and 29.6% and 28.0% in the corresponding LLM Judge scores. Its edited images achieve 5.8% higher TOPIQ-NR and HyperIQA scores than the strongest editing baselines. These results demonstrate that semantic-space evolution supports effective natural adversarial image generation both from scratch and through source-conditioned editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.