Are You Sure Your LLM Is Sure? Repairing and Auditing Confidence
Abstract
Confidence scores determine whether an LLM answer is accepted, reviewed, or regenerated, yet scalar recalibration can only correct errors that are visible through the score itself. We introduce RADAR (Repair, Adopt, and Diagnose After Recalibration), a reliability framework that separates confidence repair from deciding whether the repair should be deployed and auditing the errors that remain. Under squared loss, we decompose prediction risk into outcome noise, information hidden by the score, and mapping error; restricting repair to order-preserving maps additionally introduces a pooling gap. For repair, we propose Probe-Anchored Isotonic Calibration (PAIC), which learns a monotone map from correctness-probe predictions on unlabeled target prompts. We show that PAIC's excess risk depends on the probe's score-conditional bias and fitting error, rather than on its overall prediction accuracy. For adoption, we derive a paired procedure with finite-sample control of false improvements. For auditing, we develop a prompt-clustered kernel test that detects feature-aligned residual errors while accounting for dependent answers from the same prompt. Across multiple LLM checkpoints, benchmarks, and distribution shifts, we find that richer predictors can improve confidence estimation, but projecting their information back to a scalar score can lose these gains, and lower prediction risk does not necessarily yield better review decisions. Across 47,400 replayed raw-score allocations, paired betting adopts candidates with lower held-out test risk in 42.95% of allocations, compared with 15.09% for one-sided empirical Bernstein. Together, RADAR provides a principled workflow to repair what a scalar confidence score can express, adopt a repair only when supported by evidence, and diagnose the systematic errors that recalibration leaves behind. Code and data will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.