PoisonedEar: Knowledge Poisoning Attacks on Retrieval-Augmented Generation Multimodal Reasoning in Audio-Centric Language Models
Abstract
Recent advances in multimodal retrieval-augmented generation (RAG) allow multimodal large language models (MLLMs) to incorporate external knowledge at inference time. Although audio-centric RAG is increasingly deployed in applications such as voice assistants, meeting summarization, and financial audio analysis, its vulnerability to knowledge poisoning through audio retrieval remains largely unexplored. We present PoisonedEar, a knowledge-poisoning attack that exposes this vulnerability in audio-centric RAG systems. Under a black-box threat model, PoisonedEar induces MLLMs from diverse model families to generate attacker-controlled responses to audio queries by injecting only a small number of malicious audio–text pairs into the knowledge database. Our experiments demonstrate that this threat is both practical and severe. In decision-critical scenarios, a handful of poisoned entries can cause an MLLM to raise false alarms or suppress legitimate warnings, achieving attack success rates of up to 80%. Moreover, representative defenses provide only limited protection against the attack. These findings establish knowledge poisoning as an effective and stealthy threat to audio-centric RAG systems and highlight the need for more robust retrieval and knowledge-base safeguards. The code is available at https://github.com/PoisonedEarProject/PoisonedEar
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.