The Post-Training Tax: The Moving Target of Calibration
Abstract
Large language models typically achieve higher predictive accuracy after post-training, but a more fundamental question is whether this improvement in accuracy is accompanied by better confidence quality. To investigate this question, we examine the changes induced by post-training from two aspects of confidence: reliability and discrimination. Across 47 model–task cells comprising 16 base/tuned model pairs, 6 post-training settings, and 3 multiple-choice tasks, we find that in the 35-cell alignment cohort, post-training improves average accuracy by 1.1 percentage points while increasing the Brier reliability component by 0.047, with reliability degrading in 32 of 35 cells; discrimination remains largely stable. We refer to this phenomenon as the reliability tax of post-training. To mitigate this reliability loss, we further evaluate several confidence calibration methods and find that simple temperature scaling can substantially reduce calibration error, but its effectiveness is not fixed: when the model, task, and prompt template are held constant while the evaluation pool changes, the median temperature required for harder pools is approximately 5 times that for easier pools, and the excess ECE induced by calibration transfer is strongly correlated with the accuracy gap between pools. Further experiments show that only 20 target labels can substantially improve calibration on a new evaluation pool, reducing held-out ECE by approximately 83.4% relative to the uncalibrated state. Overall, our study reveals the reliability tax introduced by post-training and demonstrates that calibration is strongly dependent on the evaluation pool, while a small amount of target data can support effective recalibration on new evaluation pools.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.