acceptodds
Under review as a conference paper at ICLR 2027

Knowledge Capacity Beyond Independent Facts: Shared Structure, Abstraction and Frequency

Abstract

Empirical studies find that Transformers store – bits per parameter when trained on uniform random data, but this linear capacity scaling has been established only for independent facts that are spelled out token by token in the training sequences. We ask how knowledge capacity changes when data share compositional structure, when knowledge is latent and expressed through many different surface sequences, and when items occur with unequal frequencies. Using the Random Hierarchy Model, we introduce a measure of stored knowledge that credits a model for what it has learned about a latent knowledge piece rather than about any particular sequence. For verbatim memorization of structured sequences, we find a transition from learning the shared grammar to a regime in which stored information grows linearly with parameter count, with an offset reflecting incomplete grammar learning. The capacity ratio reaches – bits per parameter after epochs and scales with the total number of training exposures as . For latent knowledge, stored information still grows linearly with parameter count even though no training sequence is ever repeated, showing that linear scaling is a property of memorization itself rather than a result of repeated training data. At the same exposure budget, the capacity ratio falls from to bits per parameter as the knowledge becomes more abstract. For unequal frequencies, capacity is allocated to frequent items first, rare items are suppressed even at fixed exposure count, and total stored information departs from a single linear law. An associative memory optimized under frequency-weighted cross-entropy reproduces this suppression and traces it to interference between stored associations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.