acceptodds
Under review as a conference paper at ICLR 2027

Language Models as Information Ratchets: An Information-Thermodynamic Perspective on In-Context Learning

Abstract

In-context learning raises two questions usually studied apart: how much a prompt can teach a model, and how much task diversity pretraining needs before a model learns in context rather than memorizing. Reading a language model as an information ratchet—a machine that turns predictions into work—makes one quantity answer both: the mutual information between task and context, the learning toll. The toll is a floor no model can beat. It is tight for transformers trained on a fresh task in every sequence, so their whole loss curve follows from the task family alone; frozen language models pay several times it, and what they overpay is a failure to keep learning, not a wrong prior. The toll also counts tasks: a context tells apart only about of them, so beyond that memorizing stops paying. Trained transformers switch where it says, until the pool outgrows what they can hold, and that point is set by the task family and its prior, not by the number of a task's parameters. Post-training a pretrained model shows the same selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.