acceptodds
Under review as a conference paper at ICLR 2027

Knowing When to Defer: Training Language Models to Generate an Answer and Its Confidence Level Together for the Thresholds That Matter

Abstract

A language model that reports its confidence level in each answer allows a user to accept the answer when that confidence level meets a threshold, and to defer it otherwise. This threshold depends on the relative costs of a deferral and of a wrong acceptance. Applications usually face only a limited range of such costs, which gives a *band* of meaningful thresholds. Previous work trains models under the Brier score, which ignores this and weights every threshold equally, so some of the training signal goes to confidence values on which no decision depends. The clipped Brier score weights only the threshold band, but cannot distinguish confidence reports on the same side outside it, and thresholds that may come into use later are not weighted either. We ask how to prioritise current decisions while accounting for learning behaviour and future adaptation. Our approach is the first built to tackle this aim. We first propose a reward that combines a binary correctness score with a weighted confidence score. Our confidence score mixes the clipped Brier score with a small full-range Brier term, the *reserve*, in a controllable proportion that we call the *reserve weight*. For a fixed answer, any positive reserve weight makes the true probability of correctness the unique minimiser of the expected score, so the band concentrates the learning signal without moving its target. We then use this score for three different training systems that each allow a confidence level to be reported for a model's generated answer: a confidence reader trained on a frozen model's features; a separate model trained to report confidence for fixed answers; and a model trained to answer and report confidence jointly by reinforcement learning. Our analysis shows how the band, the reserve weight and the score's weight against the correctness reward set the incentive to answer correctly, giving a rule for placing and weighting the band. With the answer model fixed and only the confidence trained, our score gives the best-calibrated confidence among all previous methods we evaluated, and lowers expected calibration error to less than half that of a model trained under the full-range Brier reward (0.030 against 0.067 on HotpotQA, 0.034 against 0.093 on MedQA). In joint training on MedQA, our score not only cuts *decision cost* at the targeted band against the full Brier reward, but also shows better calibration and ranking over the whole range. These tests support one theory to place the band where the model's confidence lies and weight it as strongly as the incentive to answer correctly allows, which keeps the model focused on the thresholds that matter.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.