acceptodds
Under review as a conference paper at ICLR 2027

Calibrating LLMs via Region-Specific Temperature Refinement

Abstract

Large language models (LLMs) achieve strong performance across a wide range of tasks, yet post-training procedures such as instruction tuning and preference optimization often make their predicted probabilities substantially overconfident, undermining the reliability of their confidence estimates in downstream applications. We identify the reliability of the representation space as a key factor in sample-wise calibration: in reliable regions, correct and incorrect predictions are well separated, whereas in noisy regions they are mixed. Existing sample-wise methods predict temperatures uniformly across both, and therefore overfit in noisy regions and generalize poorly. To address this, we propose TS3, a theoretically grounded three-stage post-hoc framework that progressively refines temperatures from the global to the sample level. TS3 first applies global temperature scaling, then learns sample-specific temperature shifts in reliable regions while regularizing shifts toward zero in noisy regions, and finally learns a dedicated temperature for the noisy regions. Experiments on 50 tasks from three benchmarks with four instruction-tuned LLMs show that TS3 consistently achieves the lowest calibration error among a range of global, sample-wise, and prompt-based baselines. These results demonstrate the benefit of combining global, sample-level, and noise-aware calibration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.