BubbleRAG: One Question, Many Possible Paths to Evidence in Knowledge Graphs
Abstract
Graph-based retrieval-augmented generation supports knowledge-intensive question answering by connecting evidence across documents. Retrieving useful evidence, however, requires determining how query concepts are represented in the graph and how their supporting facts are connected. A query concept may match multiple nodes or edges, and the usefulness of each match can depend on its connections to other supporting facts. Pruning these alternatives too early may exclude useful evidence, while exploring all of them extensively can be costly. Moreover, relevant evidence may follow different structural patterns, making the resulting candidates difficult to compare. These issues create three challenges: preserving plausible concept matches, exploring diverse connections, and comparing candidates with different evidence structures. To address these challenges, we propose BubbleRAG, a training-free retrieval framework. BubbleRAG first extracts query concepts and retrieves candidate nodes or edges for each concept. It then constructs connected candidate subgraphs to assess these matches together with their connections, and ranks the subgraphs by query-concept coverage and semantic relevance. The top-ranked subgraphs provide starting points for LLM-guided exploration to retrieve additional facts needed for answering. Experiments on HotpotQA, MuSiQue, and 2WikiMultiHopQA show that BubbleRAG achieves the best average F1 and LLM-judged accuracy among the evaluated baselines with both 8B and 30B backbones, with the largest F1 gain reaching 7.99 points on MuSiQue.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.