acceptodds
Under review as a conference paper at ICLR 2027

MultiRAG-Bench: Benchmarking Joint Retrieval versus Knowledge Graph Fusion for Multi-Domain RAG

Abstract

Graph Retrieval-Augmented Generation (RAG) systems are widely deployed, yet a fundamental question remains unexamined: when multiple domain-specific Graph RAGs are independently constructed over the same corpus using different extraction models and prompts, should their knowledge graphs be fused into one unified structure or queried jointly while preserving domain separation? We present MultiRAG-Bench, the first benchmark for this problem. Our dataset comprises five domain-specific knowledge graphs built from a common document corpus, a multi-model heterogeneity analysis demonstrating that different LLM architectures produce substantially divergent graph structures even under identical prompts, and a stratified question set whose cross-domain requirements are each verified via direct Neo4j graph traversal, ensuring that cross-domain questions genuinely require evidence from multiple graphs. We establish baselines for both graph fusion and joint retrieval paradigms, finding that fusion incurs a hub noise cost that grows with retrieval depth, that our cache-based joint retrieval surpasses mean single-domain performance on moderate cross-domain questions, and that a substantial performance gap persists on harder multi-hop scenarios. We release all code, data, and baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.