Speculate with Memory: Lossless Acceleration for LLM Agents
Abstract
Speculative execution accelerates large language model (LLM) agents by using a smaller, cheaper model to predict future actions or observations and pre-launch subsequent work. However, many existing speculators are stateless and cannot benefit from experience accumulated across tasks. We equip the speculator with three online memory components that learn from past agent trajectories: a contrastive transition table tracking action-sequence statistics, an episodic memory retrieving contextually similar segments, and a confusion tracker suppressing recurring errors. We evaluate this approach on six benchmarks spanning three speculation types: action prediction, observation prediction, and chained prediction. Compared with stateless speculation, memory augmentation improves action-prediction accuracy by approximately 20–39% and achieves up to 2.5 times the observation-prediction accuracy on benchmarks with repetitive structure. These benefits generalize across speculator models of varying cost, with further gains from memory accumulation in domains with recurring patterns. Predictions are verified before speculative results are reused, and pre-launched work is restricted to side-effect-free operations. This preserves the actor’s trajectory while reducing latency through overlap with ongoing computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.