acceptodds
Under review as a conference paper at ICLR 2027

Learning Dynamics of LLM Calibration under Supervised Fine-Tuning

Abstract

Calibrated confidence helps assess the reliability of an LLM’s response. In the literature, LLM fine-tuning is often associated with poor calibration, but previous analyses only compared the calibration performance of pre-trained and post-trained models, leaving the training dynamics largely unexplored. In this work, we study the calibration dynamics of LLMs during standard supervised fine-tuning (SFT), and find that the calibration of an overconfident LLM is improved first and then worsens during training. The underlying reason is that hard training questions that LLMs generate incorrect responses can improve calibration, but their influence weakens as their proportion decreases during training. Based on this insight, we propose State-Gated Corrective Replay (SGCR), a novel method that amplifies the impact of those hard training questions through oversampling. In effect, SGCR can reduce calibration error and improve calibration stability during training, thereby extending the iteration window with excellent calibration. Extensive experiments on MMLU and MedMCQA validate the effectiveness of our method, reducing time-averaged ECE by 36.0–47.6% relative to standard SFT across the evaluated models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.