acceptodds
Under review as a conference paper at ICLR 2027

WARP-G: Workload-Aware Regional Provisioning for GraphRAG

Abstract

GraphRAG connects evidence across documents through entity relations and graph propagation, but constructing a full-corpus graph can incur substantial offline cost. Efficient GraphRAG methods typically simplify construction or select structurally important content. These approaches do not directly account for uneven query demand or regional differences in the benefit of graph retrieval over a passage baseline. We introduce **WARP-G** (Workload-Aware Regional Provisioning for GraphRAG), which uses an evidence-annotated design workload to select regional graph indexes. WARP-G partitions the corpus using query co-access and weak semantic links, prioritizes a limited set of graph probes by demand, baseline evidence gaps, and regional size, and measures their retrieval gains. It then selects graphs with positive conditional utility on a shared design sample and reuses their indexes at inference time, alongside full-corpus passage retrieval. We evaluate WARP-G with HippoRAG2 on four QA benchmarks. Retained indexes account for 0–14.0% of Full Graph construction tokens, and first-use graph-token cost falls by 38.7–67.2%, including design and online retrieval. On 2WikiMultihopQA, WARP-G retains one graph costing 8.2% of Full Graph construction tokens, while improving ER@10 by 1.00 percentage point and answer F1 by 1.74 points over Full Graph. WARP-G thus reduces graph-related cost while preserving or improving answer quality relative to the passage baseline on the evaluated workloads.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.