CatchUp: Replay the First, Reuse the Rest for Prefix Caching in Hybrid LLMs
Abstract
Hybrid language models can incur substantial recovery latency even when a long prefix is cached: token-indexed full-attention (FA) key–value (KV) caches can match beyond the latest recurrent checkpoint, a snapshot of recurrent-layer states at a prefix boundary. Exact recovery recomputes the full model across this gap, limiting the latency benefit of prefix reuse. Denser checkpoints reduce the gap but increase state storage. We introduce CatchUp, which loads the checkpoint and replays a recent tail only through the recurrent layers preceding the first FA layer. Unlike deeper recurrent layers, which require contextual hidden inputs from the preceding network, these layers can be replayed from token embeddings without an additional per-token input cache. A configurable fraction of the checkpoint gap controls replay work. Deeper recurrent layers retain their checkpoint states, while all FA layers reuse matched KV; the full model then processes the uncached suffix. Across three hybrid models on LongBench question answering and summarization, replaying 10% of the checkpoint gap achieves 4.32–7.94× paired time-to-first-token (TTFT) speedups over full-model recovery (Exact) at an 8K gap on 28 timing inputs. Quality differences from Exact range from −1.83 to +0.14 points on all 139 quality questions at that gap. With a 50% window, Qwen3.6-27B reaches 47.39, close to Exact at 47.44. Fixed-history tool-use evaluation on BFCL and long-context evaluation on RULER extend the study to input budgets up to 64K. On RULER at a 64K budget and a 32K gap, CatchUp differs from the full-prefix quality reference (Reference) by −2.61 to +0.20 points across the three models, compared with −4.44 to −0.07 for Tail-Replay.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.