acceptodds
Under review as a conference paper at ICLR 2027

Looped Latent Attention: Cross-Loop Key-Value Compression For Looped Transformers

Abstract

Looped, weight-tied Transformers reduce parameters by reusing a single block, but decoding still stores a separate key/value (KV) cache for every recurrence step. We show this loop-indexed cache is highly structured: for a fixed token, layer and head, the per-loop K/V vectors trace a short low-rank trajectory across loops, while the head and layer axes stay much flatter. We introduce Looped Latent Attention (), a post-training cache codec that stores compact K and V latents and reconstructs loop-specific vectors only when attention reads them. The codec is initialized from the SVD of teacher activations and refined by logit and attention-output distillation; an -2D variant folds heads into one latent for the extreme-compression regime. At matched cache budget, per-head beats head-axis MLA, cross-layer sharing, KV quantization and zero-parameter final-loop reuse; the last result shows that an unadapted endpoint cannot replace the loop trajectory. The advantage holds on Ouro-2.6B-Thinking and transfers to Huginn-3.5B, where an SVD codec is near-lossless to on decoder-independent fidelity. The same full-RoPE codecs that produce the accuracy results also serve: with a fused latent-read kernel, decodes end to end on one B200 at the teacher's fixed-batch throughput and – its peak throughput while caching – less, with decoder-independent quality unchanged. Serving through the streaming latent store exposes a train/serve mismatch, which we trace to drift of the stored K/V of generated tokens and remove by fitting the codec under the served semantics with an exact 64-token window: served GSM8K rises from to at on Ouro-1.4B (oracle re-forward ) and from to on Ouro-2.6B, matching the oracle, with the same fused kernel. An absorbed variant, LLA-D, reads latents without reconstruction and reaches the teacher's peak throughput. On student-generated prefixes, on-policy refinement raises MATH-500 at from 0.43 to 0.64–0.66 across two runs and reduces no-answer generations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.