Think Less, Retrieve Earlier: Self-Speculative Retrieval for Efficient Retrieval-Augmented Generation
Abstract
Retrieval-Augmented Generation (RAG) enhances LLM reasoning with external knowledge, but the sequential process of iterative retrieval and generation introduce substantial end-to-end latency. Existing acceleration methods primarily mitigate this overhead by limiting retrieval calls, reducing retrieval waiting, compressing retrieved contexts, or simplifying LLM reasoning. However, these methods typically reduce latency from either reasoning or retrieval perspectives, with little attention to both. We empirically observe that retrieval queries derived from intermediate-layer token reasoning are similar to those produced by full-layer LLM decoding, which is promising for invoking retrieval earlier while avoiding unnecessary full-layer computation. Inspired by this, we propose FastRAG, a new self-speculative retrieval framework that accelerates query generation through dynamic early-exit reasoning. FastRAG introduces an exit trigger to determine the optimal exit layer and a token calibrator to correct intermediate-layer token predictions, enabling earlier speculative query generation with less layer-wise computation. Next, we introduce SpecGRPO, which extends GRPO to optimize speculative-query reliability and early-exit efficiency. At inference, query entropy enables self-verification of speculative queries: low-entropy queries are directly accepted, while high-entropy queries fall back to full-layer LLM query generation, thereby reducing both retrieval waiting time and reasoning latency. Experiments on seven retrieval-augmented reasoning benchmarks demonstrate that FastRAG reduces the average end-to-end inference latency from 3.673 s to 1.382 s compared with Search-R1, while improving the average EM from 43.1% to 47.3%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.