acceptodds
Under review as a conference paper at ICLR 2027

Particle-KGRAG: Answer-Aligned Confidence for Knowledge-Graph Question Answering

Abstract

Knowledge-graph question answering requires confidence scores that distinguish correct answers from retrieval failures, yet a score used to guide graph search need not predict answer correctness. We introduce Particle-KGRAG, a training-free construction of answer-aligned confidence. It samples independent paths from a Personalized-PageRank-guided graph prior, scores completed paths with a terminal LLM judge, and aggregates normalized weights by extracted answer. Confidence is the posterior mass assigned to the answer actually returned. Using the path prior as the proposal gives a self-normalized importance-sampling formulation with a fixed budget of 129 LLM calls per query. Across six benchmarks and three seeds, with post-registration analysis changes disclosed, mean correctness AUROC is 70.6 versus 63.7 for a ToG-style beam's search-derived score; improvements on HotpotQA, 2WikiMultiHopQA, and CWQ survive paired and query-clustered testing with false-discovery-rate correction. A separately pre-registered rerun correcting prompt rendering gives 66.4 versus 58.7 and preserves these three improvements. On archived 2WikiMultiHopQA runs, selective accuracy at 50% coverage increases from .337 to .573, while the full-coverage accuracy difference is not statistically established. A beam control with answer-level aggregation reduces the AUROC gap on that dataset to 1.3 points, leaving the contribution of independent sampling unresolved. There is no demonstrated answer-accuracy gain or corrected superiority over a perturbation-based uncertainty baseline, and MetaQA-3hop favors the beam. These results characterize the utility and limits of answer-aligned confidence from judge-weighted KG paths.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.