acceptodds
Under review as a conference paper at ICLR 2027

Scalable Attention: The Inference Compute-Cache Scaling Law

Abstract

We present Scalable Attention (SA), which scales attention compute without enlarging the KV cache, adding marginal parameters. Based on SA, we propose the compute-cache scaling law, which models loss as separate power laws in compute and cache. Given unit prefill and decode costs, set by prefix caching, workload, speculative decoding, and hardware, the law returns the optimal compute and cache for any target loss. At the 3.3B scale, SA with twice the attention compute outperforms GQA in validation loss while using only of its KV cache.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.