acceptodds
Under review as a conference paper at ICLR 2027

On The Mathematical Theory Of Large Language Models

Abstract

Modern transformer-based AI methods model natural language in an explicitly statistical manner, learning a distribution over texts in training, and sampling from that distribution to perform inference. We present a general mathematical theory of such statistical language models. The theory makes no explicit use of the detailed computational structure of large language models (LLMs), treating them instead as empirical representations of the distribution over texts. Seen in this light, LLMs can be recognized as stochastic processes, indexed by the natural numbers (denoting location in a text sequence) and valued in finite sets (token vocabularies). The LLM's only responsibilities are to learn an approximation to the probability measure over text sequences, and to sample from that measure. The theory thus relies on the structure of natural language distributions to explain properties and behaviors of LLMs. We introduce the idea of natural language as a statistical mixture model, wherein documents pertaining to a specialized topic are sampled from a dialect distribution specific to that topic, and the distribution of the full training corpus is a mixture model whose components are the dialect distributions, weighted proportionally to the prevalence of their samples in the corpus. We analyze the consequences of this assumed structure for model prompting, and for fine-tuning. We confirm the theory's predictions by training a small bert model on samples generated from a small model language with a known distribution. We also check predictions with natural language, using the Library of Congress Classification (LCC) to supply dialect labels.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.