Teaching Language Models to Say What They Know: Reinforcement Learning with Internal Confidence Alignment
Abstract
As language model adoption increases, AI overconfidence can affect more people, increasing the total harm—the severity of its effects multiplied across the population exposed. Human communication relies in part on hedging; people qualify claims when they are uncertain, helping others judge how much to trust a statement. Perpetual overconfidence breeds mistrust. Language models should therefore express confidence that tracks their internal uncertainty, but standard post-training does not ask them to. We present Reinforcement Learning with Internal Confidence Alignment (RLICA) which utilizes semantic entropy estimates during the post-training process to align internal confidence of the model with a proper scoring rule tied to the probability of language model correctness. Against outcome supervision, RLICA lowers ECE from 0.221 to 0.189 while raising accuracy from .594 to .625 out-of-distribution. Its confidence report is coherent: when the model's own samples disagree it states a confidence equal to its accuracy rather than refusing to answer, and its verbalized confidence is low on 88 percent of the answers whose prose states doubt both in distribution and out, against 45 percent in distribution and 31 percent out under outcome supervision, where the verbalized confidence stays high while the prose hedges. We take this work as promising evidence for the paradigm of supervising verbalized confidence against the model's own internal confidence to produce faithful calibration of language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.