Deriving Scaling Laws for LLM Behavior via Rational Use of Representational Resources
Abstract
Scaling laws forecast language model loss from parameter count and dataset size, but they underdetermine behavior (language models at a given loss can behave arbitrarily differently), and thus cannot model why, or when, behaviors emerge in models. We propose, instead, to rationalize language models not as suboptimal solutions to an objective specified by the data, but as the optimal solutions to an objective jointly defined by the data and the resources available (parameter count, dataset size, etc.), effectively as zero-loss models of a distribution compressed by finite resources. By doing so we connect this approach to the theory of lossy compression. We design a general method to solve for a data-generating process' optimal solutions and show that, in synthetic language modeling tasks, models are generally optimal to within 0.02 nats/token. While these optima are intractable to compute exactly for large corpora, we present a scalable approximate method and analyze in-the-wild language model scientific knowledge as optimal with respect to compressed research taxonomies. Across scales, we are able to predict which facts models acquire, and, for unacquired facts, we predict the evolution of hallucinations with scale, providing empirical support for the theory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.