acceptodds
Under review as a conference paper at ICLR 2027

Think Less, Retrieve Earlier: Self-Speculative Retrieval for Efficient Retrieval-Augmented Generation

Abstract

Retrieval-Augmented Generation (RAG) enhances LLM reasoning with external knowledge, but the sequential process of iterative retrieval and generation introduce substantial end-to-end latency. Existing acceleration methods primarily mitigate this overhead by limiting retrieval calls, reducing retrieval waiting, compressing retrieved contexts, or simplifying LLM reasoning. However, these methods typically reduce latency from either reasoning or retrieval perspectives, with little attention to both. We empirically observe that retrieval queries derived from intermediate-layer token reasoning are similar to those produced by full-layer LLM decoding, which is promising for invoking retrieval earlier while avoiding unnecessary full-layer computation. Inspired by this, we propose FastRAG, a new self-speculative retrieval framework that accelerates query generation through dynamic early-exit reasoning. FastRAG introduces an exit trigger to determine the optimal exit layer and a token calibrator to correct intermediate-layer token predictions, enabling earlier speculative query generation with less layer-wise computation. Next, we introduce SpecGRPO, which extends GRPO to optimize speculative-query reliability and early-exit efficiency. At inference, query entropy enables self-verification of speculative queries: low-entropy queries are directly accepted, while high-entropy queries fall back to full-layer LLM query generation, thereby reducing both retrieval waiting time and reasoning latency. Experiments on seven retrieval-augmented reasoning benchmarks demonstrate that FastRAG reduces the average end-to-end inference latency from 3.673 s to 1.382 s compared with Search-R1, while improving the average EM from 43.1% to 47.3%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.