From Regions to Decisions: Multimodal Adversarial Attacks on Remote Sensing Vision-Language Models
Abstract
Remote sensing vision–language models (RS-VLMs) translate wide-area Earth observations into disaster-related decisions, yet their predictions may hinge on sparse local evidence and subtle variations in prompt formulation. Existing attacks typically perturb images under fixed linguistic conditions or optimize surrogate representations, overlooking the evolving interaction between visual evidence and language. We propose RS-RDA, a region-to-decision white-box multimodal attack. RS-RDA localizes decision-relevant evidence by combining proposal reliability, vision–language relevance, and semantic complementarity, and then optimizes bounded perturbations within them. After each visual round, RS-RDA compares the clean and adversarial answer distributions together with regional gradients to identify the induced score shift, the dominant incorrect competitor, and the visual regions driving the decision change. These signals guide the generation and selection of semantics-preserving adversarial prompts, enabling the prompt attack to adapt to the evolving decision boundary. Across five heterogeneous RS-VLMs, RS-RDA achieves average attack success rates of 93.89%, 90.64%, and 97.29% on building-damage counting, disaster-type recognition, and object-relation reasoning, respectively. These results exceed the strongest task-level baseline averages by 52.09%, 35.50%, and 24.00%, respectively. These results establish RS-RDA as an effective method for evaluating the adversarial robustness of RS-VLMs in disaster-scene perception and reasoning tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.