DART: Do Language Models Follow Requested Distributions?
Abstract
We ask whether language models can sample from a probability distribution that a user explicitly requests. We introduce Distribution Audit via Requested Targets (DART), an audit that asks a model for a single draw from a distribution stated in the prompt and scores the completed answer exactly over an enumerated candidate set, separately measuring whether the model attempts the task and, when it does, how closely its choice follows the requested law. Applied to 13 base and post-trained model pairs from six families, DART finds a consistent pattern: base models usually decline the task; when they answer, they typically follow the requested law more closely than their post-trained counterparts, which almost always answer yet concentrate on a single favorite outcome. Four of the six comparable chat-model endpoints place most of their probability on a single answer in a uniform ten-way draw, and the favorite shifts with the answer range rather than a fixed token. A factorial experiment shows that the residual distortion decomposes into interpretable priors over label identity and display position that predict held-out combinations but do not transfer across models or prompt templates. Interventions that flatten the measured output distribution fail to restore genuine sampling behavior. DART thus makes distribution-following measurable: current post-training trades it for unconditional compliance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.