acceptodds
Under review as a conference paper at ICLR 2027

Probabilistic Calibration is a Trainable Capability in Language Models

Abstract

Language models are increasingly used in settings where outputs should follow an intended distribution, yet their generation probabilities are often poorly calibrated to it. In many such requests the distribution is implicit and the valid outputs cannot be listed in advance, so the task cannot be passed to an external random number generator. We study whether this capability can be improved directly through fine-tuning. Concretely, we fine-tune language models on synthetic prompts that require sampling from mathematical distributions, and compare two Calibration Fine-Tuning variants: a soft-target method that converts the desired output distribution into trie-derived next-token targets, and a hard-target method that trains on sampled completions from the same target distribution. Across 12 models spanning four families, both methods substantially improve structured-sampling fidelity on held-out distribution families and unseen parameter settings, showing that training on cheap synthetic mathematical distributions improves sampling on unseen distributions, even in natural language. The gains also transfer to broader stochastic generation benchmarks, including open-ended random generation, where fine-tuned models broaden their first-token support by one to two orders of magnitude, and NoveltyBench. The gains sometimes reduce downstream capability, especially arithmetic reasoning, with costs varying by model. Overall, our results show that probabilistic calibration can be improved through fine-tuning with either objective, and that calibration learned on mathematical distributions transfers to natural-language stochastic generation. Code is available at https://anonymous.4open.science/r/calibration-finetuning-3C11.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.