Decide, Don't Decode: Turning Post-Training Language Models into Calibrated, Coherent Decision Models
Abstract
Many systems use language models for bounded decisions, such as routing a ticket or flagging a message, and act on the model's confidence. We develop a post-training method that turns an open generative model into a typed decision model: for a choice, a rating or a yes/no question it returns probabilities over the allowed answers without decoding text, trained to be calibrated and coherent across related questions. We build DECIDE-22, 22 typed decision tasks with 22,773 test items; on it, decoding the answer with a stated confidence is no more accurate than reading these probabilities, costs several times more, and yields a confidence that barely separates right from wrong answers (AUROC 0.59, and 0.69 after reasoning, against 0.76). Tracing 26 checkpoints of six model families through every post-training stage they publish, we find that preference optimization sharpens the supervised readout into over-confidence at unchanged accuracy, almost exactly as a temperature would ( 0.92–0.95 in three pipelines). Fine-tuning the readout with a proper scoring rule reverses this sharpening and improves accuracy and calibration, but increases contradictions across logically related questions. We address this with a penalty based on the de Finetti sure loss of each question and its automatically derived siblings. The penalty needs no additional labels and reduces incoherence tenfold without a measurable loss in accuracy. A 9B open model trained this way reaches 76.0% accuracy and a calibration error of 0.054, against 73.3% and 0.113 for a commercial decision API as served (0.068 when given the same calibration recipe), with under a third of its incoherence; at 4B it also beats a concurrent open decision model trained on six times the data. The API keeps its lead on the six held-out tasks, most of it from a single rating scale. We release the benchmark, code and trained models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.