Where to Look, What to Reveal: Privacy Preserving Multimodal RAG with Multiple Agents
Abstract
Modern large language model (LLM) systems increasingly rely on centrally incorporating task-relevant external context at inference time, in addition to scaling model capacity. Retrieval-augmented generation (RAG) provides a general mechanism for such context augmentation, yet existing cloud-centric RAG systems often require fine-grained user data to be centrally accessible, creating unnecessary privacy exposure when personalized information is distributed across heterogeneous local sources. We propose PPM-RAG, a model-frozen multi-agent multimodal RAG framework that enables collaborative retrieval over distributed private data while keeping raw documents local. At its core, PPM-RAG introduces a central–local Federated Schema Directory (FSD), where local agent specialists organize modality-specific documents into hierarchical semantic directories and expose only compact root-level summaries and metadata to the central agent. During inference, the central agent decomposes a user query into sub-queries and traverses the central FSD to identify relevant specialists. Selected specialists then perform hierarchical retrieval within their local FSDs and return only query-relevant evidence for final evidence-grounded reasoning. This design separates the discovery of where relevant knowledge resides from the disclosure of what information is required, enabling efficient and privacy-aware collaboration across distributed multimodal sources. Extensive experiments on WebQA and MMDocRAG against six baselines show that PPM-RAG improves retrieval Recall by 12.4% on average and answer quality by up to 32.9%, while reducing retrieval time by 90.6% compared with OpenRAG. PPM-RAG further reduces IIR by 10.8% and improves UMBRELA by 28.1%, demonstrating effective, efficient, and privacy-aware multimodal retrieval.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.