Enriching Document Knowledge Bases through Context Propagation
Abstract
Understanding real-world document collections, such as conversations or codebases, requires more than retrieving isolated facts: it involves reconciling information scattered across documents and forming a coherent view of entities and events. This can be challenging because authors often assume shared context, using *implicit cross-document references* whose interpretation depends on information in other documents. To isolate this challenge, we introduce XRefQA, a diagnostic benchmark of small synthetic document collections in which answering questions requires resolving these references and combining evidence from multiple documents. Even Claude Opus 4.6 achieves only 7% accuracy on XRefQA using standard retrieval-augmented generation (RAG). To address this limitation, we propose *context propagation*, which iteratively enriches each document with source-grounded annotations using evidence retrieved from related documents. Enriched documents inform subsequent rounds, allowing resolved references and related facts to propagate across the collection. After two propagation rounds, the weaker Claude Sonnet 4.6 achieves 80% accuracy with RAG over enriched documents. Beyond implicit-reference resolution, the method also improves multi-hop question answering on MuSiQue, 2WikiMultiHopQA, and HotpotQA, as well as fact discovery and source attribution on SummHay. These results suggest that making cross-document connections explicit during preprocessing helps downstream systems retrieve, combine, and attribute distributed evidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.