How Faithfully do Language Models Acquire Probabilistic Knowledge?
Abstract
Many high-stakes decisions require calibrated distributions over plausible outcomes, not just top-ranked answers. LLMs are increasingly used to support probabilistic reasoning, yet it remains unclear how faithfully they acquire probabilistic knowledge during pretraining. We study this question in a controlled setting by generating synthetic Bayesian networks with known ground-truth distributions, verbalizing samples from these networks into text, and continually pretraining on the resulting corpora. We measure how well the models recover the underlying conditional distributions and the relations among them while varying network size, documentation policy, model scale, and training exposure. Models recover individual conditional probabilities closely, but the relationships among them exhibit what we term relational compression: models largely preserve which outcome is more probable and in which direction evidence moves a prediction, but they systematically underestimate the difference in likelihood between pairs of outcomes and how strongly evidence should shift a prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.