acceptodds
Under review as a conference paper at ICLR 2027

Scaling Laws for LLM Memory Sparsity

Abstract

Sparse memory architectures offer a route to expanding language-model capacity while accessing only a small fraction of parameters per token. Their finer access granularity promises greater sparsity than conventional mixture-of-experts (MoE), but how memory capacity should scale with active computation and training data remains insufficiently understood. We introduce and validate scaling laws for Product Key Memory (PKM) that capture how active parameters, memory capacity, and training tokens jointly determine loss. Our analysis spans 421 measured training endpoints from 44 configurations of dense Transformers, PKM, STEM, and MoE, with total parameter capacities reaching 21.9B and training budgets reaching 128B tokens. A joint PKMMoE fit with shared coefficients worsens held-out MoE predictions, motivating separate scaling relationships for the two architectures. Using these architecture-specific fits, we find that optimal memory allocation depends on the balance between capacity and compute: increasing capacity at fixed compute favors greater sparsity, whereas increasing compute at fixed capacity favors more active parameters. Among the studied variants, predicted loss favors PKM when capacity is abundant relative to compute, MoE in an intermediate regime, and dense Transformers when compute is abundant relative to capacity. A complementary roofline analysis estimates a PKM throughput advantage over MoE in the modeled mobile inference setting, where selective retrieval reduces flash-memory traffic. Together, these findings identify when sparse memory offers advantages in predictive performance and execution efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.