AdaMM-RAG: Adaptive Multi-Source Multi-Modal RAG via Planner-Guided Coordination
Abstract
RAG grounds LLMs in external knowledge, but existing multimodal multi-agent systems usually use a fixed topology that calls all sources in parallel and aligns only vision-language evidence. This causes two issues: irrelevant sources add noise and cost, and cross-source evidence is never unified pre-fusion. Adaptive agent methods remain text-only, and do not route across sources or modalities. We propose AdaMM-RAG, which enables source-level adaptive routing via a Planner Agent that selects per-query which sources to consult. A Cross Alignment Agent unifies selected heterogeneous evidence; a Decision Agent refines and verifies answers with CoT; a quality gate triggers fallback only when evidence is insufficient. On ScienceQA, AdaMM-RAG achieves 95.51% accuracy, surpassing HM-RAG (SOTA multi-agent RAG) by +1.78% and MAO-ARAG by +12.02%; on CrisisMMD, it reaches 80.20% accuracy, outperforming HM-RAG by +21.65% and MAO-ARAG by +5.77%, while reducing retrieval cost by 55.5%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.