acceptodds
Under review as a conference paper at ICLR 2027

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

Abstract

Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase-sensitivity. In the DeepSeek-V4 family, long context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase-sensitivity across the tested variants. Mechanistic analysis using causal interventions in these models and DeepSeek-V4 reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression therefore requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.