acceptodds
Under review as a conference paper at ICLR 2027

Foldable Maps Enable Expressive and Efficient Attention

Abstract

Long-context inference is bound by the KV cache. Linear attention on its own struggles with precise recall while dense attention has higher precision but becomes extremely expensive for long sequences. Shared latent compression (MLA) improves decode bandwidth by reading once across all heads, but that leaves no benefit for sparse head access. In this paper we introduce Foldable Maps, a conditional linear map system which pairs a shared latent with per-expert latents that utilize a foldable workspace. The expert-level latent expands using 6 of 24 map projections per-expert per-token, which simplifies to one once selected. This enables the training-time workspace to fold with absorption, making the per-expert cost independent of workspace and head width for inference. At 815M parameters with token-matched arms and a 4.57x KV compression ratio over its parent model, this format matches or exceeds both uncompressed baselines on NIAH retrieval averaged over 4k-100k contexts (+20% over parent) and more than doubles the score of the next best measured compression format at equal latent width. The perplexity is the best of the measured compressed formats, but is still  5% worse than the uncompressed parent at this compression ratio. We also train a 24-byte per-token per-layer indexer which maintains the full NIAH retrieval edge (0.423 vs 0.426) with  12x less read traffic compared to the parent model at 100k context length. Partitioning by consumer makes expert count a free axis for attention capacity, aligning the training and serving benefits of attention-side MoE.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.