Rewind: Restoring Evicted Prefix States for Hybrid LLM Agents
Abstract
Long-context LLM agents repeatedly revisit shared histories during retries and branch exploration, making efficient prefix reuse essential for responsive serving. However, hybrid models cannot reconstruct an earlier recurrent state from the latest one, so cache eviction forces requests to recompute long prefixes even when they add only a short suffix. Optimizing checkpoint placement alone cannot recover the state once it has been discarded. Our key insight is that preserving complete boundary checkpoints can decouple prefix reuse from the lifetime of native cache entries. We present Rewind, a serving system that restores evicted prefixes for hybrid LLM agents. First, Rewind persists recurrent state, convolution history, and attention keys and values at a common token boundary, coordinating their installation across workers before execution resumes. Second, it uses cost-based checkpoint admission and retention to prioritize expected saved computation under storage pressure. Finally, its restore gate compares complete installation cost with the additional prefill avoided beyond the native cache hit. On a 16-session Qwen3.5-27B-FP8 replay using two RTX 4090 GPUs, Rewind’s additional checkpoint store reduces prefill by 91.7%, replay duration by 30.5%, and return-round p95 latency by 26.5% relative to native caching. Under forced eviction at matched storage capacity, selective Rewind performs 41% less prefill than Snapshot-LRU, which saves every boundary and evicts by recency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.