Calibrating Fine-Tuned Models via Selective Confidence Reuse
Abstract
Fine-tuning can improve a model's predictions without making its confidence a reliable guide to whether those predictions are correct. Calibration based only on the final confidence score cannot distinguish equally confident predictions with different probabilities of being correct, even when the model before fine-tuning retains information that separates them. We recast calibration as selective inheritance: preserve reliable information from the model before fine-tuning (the predecessor) where confidence remains stable, and repair confidence estimates elsewhere. Our method, A-PAIR, aligns predecessor scores into a reference and uses labeled source data and unlabeled target inputs to audit where this reference can be retained and where correction is needed. Our theory shows that predecessor information can recover distinctions inaccessible to score-only calibration, and that selective inheritance can reduce finite-sample estimation error when repair is localized. On public benchmarks under controlled distribution shifts, A-PAIR improves the calibration of fine-tuned language models while keeping their predictions accurate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.