acceptodds
Under review as a conference paper at ICLR 2027

The COUNT: Counting N-grams Unlocks Scaling on Host Memory and Training-Free Adaptation

Abstract

The scaling relationship of compute and performance for language models (LM) has been extensively studied; more recently, models are being designed to scale memory use, as well. In particular, adding input embeddings for n-grams scales memory separately from compute use, because only individual rows from the embedding table are read during each pass. However, existing methods train embeddings from scratch, which needs expensive GPU memory for fast reading and writing. Meanwhile, the ordinary memory of the host system is cheap and underutilized by these training methods. To scale on host memory, we introduce COUNT (Counting Offline Unlocks N-gram Transformers). COUNT scales n-gram representations while adding zero learned parameters by storing precomputed n-gram counts on host. We use the corpus frequencies to weight together the LM's own unembeddings of the next words most often following the current n-gram history. At a fixed compute budget, COUNT can be added as an additional input embedding for pretraining 1B-parameter LMs to improve validation loss as much as an LM with 58% more compute and task loss as much as 24% more compute. If GPU memory is limited, using COUNT to let backbone parameters fill GPU memory gives better performance than splitting memory with additional learned embeddings. Beyond efficiency, COUNT offers a new affordance: switching in a different table of n-gram counts from a domain-specific corpus at inference time to adapt to a new task without any training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.