When Weather Misleads Retrieval: Atmospheric Attacks on Remote Sensing Vision-Language RAG
Abstract
Multimodal RAG systems rely on vision-language retrievers to ground visual queries in external evidence, yet the retrieval stage itself remains underexplored as an adversarial target in remote sensing. We introduce an atmospheric retrieval attack that modifies only the input image, leaving the retriever, generator, and knowledge base unchanged. The attack overlays parameterized cloud- and haze-like patterns and optimizes them to attract image representations toward target atmospheric evidence while suppressing source-scene evidence, with additional constraints on rank separation and visual naturalness. We evaluate the attack on a seven-dataset remote sensing RAG benchmark with five CLIP-style retrievers and downstream vision-language generators. Across retrievers, optimized atmospheric perturbations consistently outperform clean inputs and multiple handcrafted, random, and fixed baselines in promoting weather-related evidence into top-ranked retrievals. On GeoRSCLIP ViT-B/32, Weather@5 rises from 0.71% to 43.29%. The resulting retrieval corruption further induces measurable weather hallucinations and semantic shifts in downstream generation. These results reveal retrieval-stage evidence grounding as a distinct vulnerability of remote sensing vision-language RAG.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.