Understand Once, Retrieve Many Times: SAADI for Single-Pass Retrieval of Distributed Evidence in Long Documents
Abstract
Long-document retrieval is challenging when relevant evidence is spread across distant passages or multiple hops, especially under a limited retrieval budget. Single-pass top-k retrievers have predictable compute and latency at query time, whereas agentic systems pursue higher recall by reading a document over several steps. We move progressive document understanding to index time, where a single-pass top- retriever can use it without query-time large language model (LLM) calls or adaptive trajectories. SAADI (Stateful Accumulation of Document Information) builds a query-agnostic document state at index time with one state-update call per chunk. Each call extracts the chunk's entities, conditioned on records recalled from the entity memory and on a hierarchical summary of the prefix. The chunk is linked to the resulting records and to the summary covering it, and the original chunks remain the evidence and retrieval units. At query time, two auxiliary scores are computed from a chunk's linked records and its covering summary and fused with any sparse or dense base retriever's score in a single top-k pass. Index-time LLM calls, including summary folds, are fixed before any query arrives, at most 1.25 per chunk in the default configuration. Each call reads a context that grows only logarithmically with the document, and appending chunks extends the state without rebuilding it. Across four datasets and five retrievers, SAADI improves HR@10 over its base retriever in all twenty pairs, by 0.3 to 12.4 points (1.1 to 4.0 under BM25). It ranks first or second in 34 of 40 dataset, retriever, and depth cells and, at matched evidence exposure, is ahead of agentic retrievers in sixteen of twenty pairs. The gains persist to 256K-token contexts and at every evidence-hop count. Agentic costs recur with every query. After a break-even number of queries per document, which we report in model calls and in tokens, they exceed SAADI's one-time construction cost. Ablations show that the entity-memory and summary signals each improve retrieval, and that fusing both is best in every setting. Thus document-level connections behind distributed and multi-hop evidence can be established once at index time, at a cost known before the first query.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.