acceptodds
Under review as a conference paper at ICLR 2027

WEAVE: Compressing State Changes Along Native Evolution in Linear Attention LLMs

Abstract

A new generation of hybrid linear-attention models replaces much of the growing KV cache with fixed-size recurrent states. Yet agent serving must retain many of these states as checkpoints for prefix reuse across turns and branches, creating substantial GPU-memory pressure and costly prefix recomputation under load. We present WEAVE, a training-free approach that encodes state changes from a shared anchor and efficiently constructs these compact representations from the model’s native updates. WEAVE combines shared-anchor residuals, recurrence-driven encoding from native model updates, and periodic compaction with fused GPU kernels for efficient online serving. We evaluate WEAVE on Qwen3.5-35B-A3B and Kimi-Linear-48B across reasoning, long-context, and agentic workloads. WEAVE compresses recurrent checkpoints by up to while maintaining task quality and outperforming INT8 and INT4 quantization on all six benchmarks. Integrated into vLLM's native prefix-cache path, the reduced state footprint lowers prefix eviction and recomputation under memory pressure, yielding up to lower TTFT and higher token goodput under SLO constraints.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.