WEAVE: Compressing State Changes Along Native Evolution in Linear Attention LLMs
Abstract
A new generation of hybrid linear-attention models replaces much of the growing KV cache with fixed-size recurrent states. Yet agent serving must retain many of these states as checkpoints for prefix reuse across turns and branches, creating substantial GPU-memory pressure and costly prefix recomputation under load. We present WEAVE, a training-free approach that encodes state changes from a shared anchor and efficiently constructs these compact representations from the model’s native updates. WEAVE combines shared-anchor residuals, recurrence-driven encoding from native model updates, and periodic compaction with fused GPU kernels for efficient online serving. We evaluate WEAVE on Qwen3.5-35B-A3B and Kimi-Linear-48B across reasoning, long-context, and agentic workloads. WEAVE compresses recurrent checkpoints by up to while maintaining task quality and outperforming INT8 and INT4 quantization on all six benchmarks. Integrated into vLLM's native prefix-cache path, the reduced state footprint lowers prefix eviction and recomputation under memory pressure, yielding up to lower TTFT and higher token goodput under SLO constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.