acceptodds
Under review as a conference paper at ICLR 2027

Calibrating LLMs by Reducing Overconfidence via Instance-Adaptive Smoothing

Abstract

Large language models (LLMs) play a growing role in real world decisions, emphasizing a need to ensure accurate reflection of their uncertainty to minimize risk. While the uncertainty calibration of neural networks precedes LLMs, the uniqueness of the setting leads to observations that cannot directly be addressed through existing methods, such as the emergence of severe overconfidence past the base LLM pre-training phase. While label smoothing (LS) have been suggested as a possible remedy due to its ability to mitigate overconfidence, approaches targeted to traditional machine learning settings can be ill-suited to the demands of LLMs, for example multiple training and alignment stages. We identify why this may fail to consider specific nuances unique to LLMs, in particular how the decomposition of the traditional label smoothing loss fails to eliminate overconfidence during misclassifications, leading to possible degradations in the generative abilities of models that are essential for further alignment and tuning. To this end, we introduce _**H**olistic **S**moothing **Hi**nges_ (___HoSHi___), an instance-adaptive smoothing approach during supervised fine-tuning to enable calibration while retaining strong generative capability. ___HoSHi___ acts as an adjustment to the learning objective that can mitigate previous concerns and enable the stable training of language models that remain calibrated at each stage. Results after different stages demonstrates the effectiveness of our approach, highlighting possible avenues towards better uncertainty-aware LLMs that are more interpretable and effective for public use.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.