The Illusion of Global Calibration: Group-Aware Risk in Class-Incremental Learning
Abstract
Reliable class-incremental learning requires confidence estimates that remain informative after sequential training, because downstream decisions such as selective prediction and human referral depend on per-sample confidence rather than on accuracy alone. Calibration is typically summarized by aggregate scores, yet such averages can hide substantial errors within groups of classes introduced at different times during the training schedule. In this paper, we show that global temperature scaling can improve overall calibration while leaving large errors in specific schedule-age groups, and that the direction of these group-level changes differs across learners. Aggregate improvement therefore does not by itself certify uniformly reliable confidence. To make this discrepancy visible, we introduce a group-aware audit that evaluates aggregate and group-level calibration jointly, together with the direction of confidence bias, the source of calibration data, and the scale of the scores being calibrated. Across class-incremental benchmarks, the discrepancy persists even in high-accuracy models, although its magnitude depends on both the temperature search range and the data used for fitting. These findings support reporting group-level calibration alongside aggregate scores whenever continual learners are compared or deployed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.