acceptodds
Under review as a conference paper at ICLR 2027

AdaMM-RAG: Adaptive Multi-Source Multi-Modal RAG via Planner-Guided Coordination

Abstract

RAG grounds LLMs in external knowledge, but existing multimodal multi-agent systems usually use a fixed topology that calls all sources in parallel and aligns only vision-language evidence. This causes two issues: irrelevant sources add noise and cost, and cross-source evidence is never unified pre-fusion. Adaptive agent methods remain text-only, and do not route across sources or modalities. We propose AdaMM-RAG, which enables source-level adaptive routing via a Planner Agent that selects per-query which sources to consult. A Cross Alignment Agent unifies selected heterogeneous evidence; a Decision Agent refines and verifies answers with CoT; a quality gate triggers fallback only when evidence is insufficient. On ScienceQA, AdaMM-RAG achieves 95.51% accuracy, surpassing HM-RAG (SOTA multi-agent RAG) by +1.78% and MAO-ARAG by +12.02%; on CrisisMMD, it reaches 80.20% accuracy, outperforming HM-RAG by +21.65% and MAO-ARAG by +5.77%, while reducing retrieval cost by 55.5%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.