acceptodds
Under review as a conference paper at ICLR 2027

GaugeRoPE: RoPE-Compatible Executable Latent Attention

Abstract

Latent KV cache compression reduces memory footprint by mapping historical keys and values into compact representations. However, generic compressed key states require reconstruction to native dimension before applying RoPE and computing attention scores, adding sequence-scaling overhead to the decode path. We characterize when a latent KV representation can serve directly as the execution state of pretrained RoPE. For RoPE-based transformers, exact static linear execution of the rotary key path requires the latent-to-native map to preserve the pretrained RoPE action through an intertwining relation, restricting exact mixing to coordinates that share the same RoPE frequency. Guided by this characterization, we build GaugeRoPE, a training-free executable latent KV representation. It combines a RoPE-compatible carrier that preserves the rotary action with an approximate position-independent residual key–value (RKV) correction for complementary key and value information, both evaluated without per-token dense historical K/V reconstruction. Experiments on Mistral-7B show competitive quality at KV-state reduction and substantially stronger RULER and LongBench performance as the state budget tightens. Additional MHA results validate the construction beyond GQA. By enabling direct latent execution, GaugeRoPE achieves the decode throughput of the vanilla uncompressed baseline on 32K-context generation with batch size 4 while halving the KV cache.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.