acceptodds
Under review as a conference paper at ICLR 2027

Rewind: Restoring Evicted Prefix States for Hybrid LLM Agents

Abstract

Long-context LLM agents repeatedly revisit shared histories during retries and branch exploration, making efficient prefix reuse essential for responsive serving. However, hybrid models cannot reconstruct an earlier recurrent state from the latest one, so cache eviction forces requests to recompute long prefixes even when they add only a short suffix. Optimizing checkpoint placement alone cannot recover the state once it has been discarded. Our key insight is that preserving complete boundary checkpoints can decouple prefix reuse from the lifetime of native cache entries. We present Rewind, a serving system that restores evicted prefixes for hybrid LLM agents. First, Rewind persists recurrent state, convolution history, and attention keys and values at a common token boundary, coordinating their installation across workers before execution resumes. Second, it uses cost-based checkpoint admission and retention to prioritize expected saved computation under storage pressure. Finally, its restore gate compares complete installation cost with the additional prefill avoided beyond the native cache hit. On a 16-session Qwen3.5-27B-FP8 replay using two RTX 4090 GPUs, Rewind’s additional checkpoint store reduces prefill by 91.7%, replay duration by 30.5%, and return-round p95 latency by 26.5% relative to native caching. Under forced eviction at matched storage capacity, selective Rewind performs 41% less prefill than Snapshot-LRU, which saves every boundary and evicts by recency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.