acceptodds
Under review as a conference paper at ICLR 2027

Compute-Optimal Agentic RAG

Abstract

Long-context language models can read whole documents, but paying for every token on every question is costly. An agent can instead bring document information into its context through a question-independent summary, written once and reused, or through retrieval at question time. How should inference compute be split between summary length, retrieval and model size? We vary five levers—summary length, number of retrieval calls, passages per call, passage length and model size (Qwen3.5, 0.8B to 27B parameters)—on LongBench-v2, over about 860k question–configuration pairs, and count each question's theoretical FLOPs exactly. Summaries and retrieval supply one resource, document information, and substitute for each other: the longer the summary, the less retrieval pays. Model size is a separate axis, with a threshold near 4B parameters below which the levers help little. A 15-parameter law built on this structure predicts lever combinations it was not fit on close to the limit set by sampling noise. Its compute-optimal allocation is simple: past the threshold, and for documents queried often enough to amortize the summary, spend about half of each question's compute on the summary and keep retrieval to a few thousand tokens; for a document queried once, retrieval and a larger model are best. An agent left to decide how often to search stays about 7 accuracy points below this allocation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.