SCAMP: Black-Box Knowledge Poisoning Attack on Multimodal Retrieval-Augmented Generation
Abstract
Large vision–language models benefit from multimodal retrieval-augmented generation, which grounds responses in external image–text evidence to mitigate hallucination. However, reliance on external knowledge bases introduces a critical vulnerability: knowledge poisoning. Existing poisoning methods are often impractical, as they rely on white-box access, target limited sample-wise scenarios, and neglect reranking components. Additionally, they underutilize the visual modality's potential and suffer from cross-modal inconsistencies. We propose SCAMP, a black-box framework that iteratively crafts aligned poisoned pairs through a self-competitive agent loop. SCAMP integrates a Visual Attacker and a Text Attacker to jointly optimize, Competing Agents to simulate benign samples and improve attack generalization, and three Evaluators to guide optimization via retrieval, generation, and consistency feedback. For class-wise attacks, we introduce agent-driven prototype synthesis to construct effective visual anchors without gradients. Extensive experiments show SCAMP consistently outperforms baselines in black-box settings and remains robust against diverse defenses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.