acceptodds
Under review as a conference paper at ICLR 2027

Depth-Aware KV Cache Compression in Looped Transformers

Abstract

Looped transformers have advantages over their counterparts by reusing the same layer parameters across recurrences. Yet they typically store separate state caches for each recurrence step. The naive KV caching scheme would incur linearly growing overhead in terms of the number of recurrences, limiting the efficiency gain from looped transformers. There exist several techniques for mitigating this overhead by layer-wise mixed-precision KV quantization according to sensitivity, as well as predicting/reconstructing the states at runtime. Our investigation tries to discover connections between these two approaches through depth analysis. As observed, replacing cached states from earlier recurrences with the last recurrence's state causes less distortion in deep layers than in shallow layers. Contrarily, retaining all states with residual coding while increasing the precision of the first recurrence and reducing that of the last causes more distortion in deep layers than in shallow layers. Motivated by this finding, we introduce a depth-aware allocator that simultaneously selects which recurrent states to retain, along with quantization and coding schemes under certain capacity constraints. The allocator was calibrated on 16 documents and tested on 256 held-out documents for the models Ouro-1.4B and Ouro-2.6B. On the test set above, our allocator decreased KL divergence by 36% and 24%, respectively, compared with layer-wise mixed-precision allocation under the same tight storage budget. Furthermore, without recalibrating, allocations calibrated using 384-token prompts still reduced KL divergence by 57% and 33% w.r.t. the layer-wise allocation scheme on separate test documents with longer prompts of 2,048 tokens.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.