acceptodds
Under review as a conference paper at ICLR 2027

CoverKG: Coverage and Compactness Dual Rewards for Knowledge Graph Extraction

Abstract

Knowledge-graph-based retrieval-augmented generation relies on an LLM extractor to build a graph index, the quality of which directly governs retrieval effectiveness and storage efficiency. Yet existing methods supervise the extractor only through downstream QA accuracy. This sparse, end-to-end signal cannot diagnose whether failures arise from insufficient semantic coverage or from redundant extractions that introduce noise into retrieval. We introduce CoverKG, a framework that characterizes index quality through two complementary objectives, namely coverage and compactness, and converts this characterization into a construction-time training signal that can be computed for every source chunk and used to optimize the extractor with reinforcement learning. Coverage aligns the source sentences with the extracted graph elements to pinpoint missing semantics at the sentence level. Compactness evaluates each element's marginal contribution relative to the other elements in the same extraction, directly penalizing redundancy. A composite reward balancing both objectives drives policy optimization. The framework is readily applicable to the graph construction stage of various graph-based RAG systems. Experiments across multiple graph-based RAG pipelines and QA benchmarks show that, even under the same downstream evaluation protocol as prior work, extractors trained with CoverKG consistently achieve higher accuracy with more compact graph indices, demonstrating that per-chunk, dual-objective supervision at index construction time yields both more effective retrieval and more efficient storage. Our code is available at https://anonymous.4open.science/r/anonymous_repo-1357.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.