Continual Calibration: Stale Conformal Thresholds Fail Under Lifelong LLM Fine-Tuning
Abstract
Continual fine-tuning changes not only a language model's predictions but also the score distribution used by post-hoc uncertainty procedures, and this creates a specific failure mode for conformal prediction: thresholds calibrated before later fine-tuning can lose task-level coverage even when top-1 accuracy changes only modestly. We measure this failure mode on sequentially fine-tuned LLMs across five model configurations and eight task sequences, seven of which are classification-style or multiple-choice settings covered by our formal framework. In the classification-style grid, stale-threshold coverage loss is substantially larger than accuracy loss; by aggregate loss, coverage degrades by roughly as much as accuracy, and the effect is diagnosed by drift in the nonconformity-score CDF rather than by argmax-accuracy drift. We propose task-specific calibration replay, which maintains a small held-out calibration buffer per task and recomputes that task's split-conformal threshold under the current model after each update. Under standard split-conformal exchangeability assumptions, and assuming the calibration buffer is not used for training or model selection, this restores finite-sample marginal coverage for the corresponding task without adding gradient updates. Empirically, using a default calibration-buffer cap of examples per task, with for low-resource tasks, task-specific recalibration (which acts as a continual Mondrian conformal predictor) reduces coverage loss to near nominal while preserving utility, with average set sizes and singleton rates close to a large fresh-calibration reference. The results hold against competitive baselines including online adaptive conformal inference (ACI) and weighted conformal prediction, both of which recover only part of the coverage gap, and a simple zero-shot router retains most of the benefit when task identity is unknown. Our formal guarantees apply to classification-style tasks with known task identity and task-specific buffers; open-ended generation results are visually separated and reported only as exploratory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.