How Much Information Can We Fit in a Transformer?
Abstract
How efficiently can transformers store factual knowledge? The answer to this question has implications for both model design and understanding the limits of transformer-based systems. Prior studies observe a linear relationship between the number of facts memorized and the number of model parameters required; substantially worse than what theoretical analyses predict is achievable. We re-examine empirical scaling: through rigorous and systematic experiments, including extensive hyperparameter search via Bayesian optimization, we determine the smallest model that can perfectly memorize a given set of facts, and measure how this minimum size scales with the number of facts. On synthetic datasets of random, independent (subject, relation, object) triples—e.g., (Paris, capital_of, France)—drawn at random with no exploitable structure, we find that memorization capacity, the size of the largest dataset a model memorizes completely, grows exponentially with the transformer dimension: the minimum transformer dimension required to memorize facts grows logarithmically in , and across 11 dataset sizes spanning to a lack-of-fit F-test does not reject the logarithmic fit (with an offset) but it does reject all other tested alternatives including power-law. Along the depth axis, however, at fixed width, the maximum memorizable dataset size grows sublinearly with the number of layers over the measured range. We further extend our exploration to structured data containing symmetric relations or two-hop composition. In this setting, we find that transformers exploit these regularities to further improve capacity. All code is available to reproduce the results at [redacted].
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.