acceptodds
Under review as a conference paper at ICLR 2027

STAPLE: Static Sparse Upcycling of Pretrained Checkpoints using Embedding Modules

Abstract

Sparse computation can scale the parametric capacity of large language models without proportionally increasing active computation, but pretraining sparse ar- chitectures remains expensive and data-hungry. Existing sparse upcycling meth- ods mainly target Mixture-of-Experts (MoE) models, where expert homogeneity, sample inefficiency, and system overheads can limit practical scalability. Static sparse memory can address the efficiency and expert-homogeneity challenges, but naively integrating it into pretrained computation yields poor returns. In this work, we introduce STAPLE, a training recipe and architecture design for successfully integrating static sparse memory into a pretrained model. STAPLE preserves the pretrained model’s computation and adds complementary token-specific knowl- edge in a scalable manner. On OLMo-2 1B, STAPLE improves over continued dense training by up to 2–3 points across math, code, discrete reasoning, and knowledge-intensive tasks. Average performance increases with static capacity, and improvements persist on the larger OLMo-2 7B backbone. STAPLE also outperforms static sparse baselines such as STEM and Engram, and achieves a higher average score than an MoE upcycling baseline while decoding 2× faster with less than half its GPU memory. Importantly, despite introducing billions of parameters late in training, our host-offloaded implementation keeps generation throughput within 1.5% and training throughput within 5% of the dense base- line, using a 1 GB GPU cache for inference. These results show that static sparse memory can provide a sample-efficient, scalable, and system-efficient path for expanding pretrained models without sparse pretraining from scratch.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.