Debiased Generalized Calibration Moments
Abstract
Modern neural networks can be accurate but poorly calibrated. We revisit train-time multiclass calibration through a generalized moment view that unifies binning, mean-matching, and kernel objectives as fixed witness function choices on the probability simplex, and we ask whether the witness itself can be learned. We propose Debiased Generalized Calibration Moments (D-GCM), a lightweight min-max regularizer that trains a small critic on predicted probabilities to expose miscalibration patterns. The key challenge is empirical estimation: the naive squared batch moment is a biased V-statistic whose diagonal self-interaction term biases the objective for any fixed witness and becomes an optimization shortcut for a learned critic, rewarding sample-wise residual energy even when the target calibration moment is zero. D-GCM removes this shortcut with the corresponding U-statistic over distinct sample pairs. We theoretically show that rich witness families characterize strong multiclass calibration, while restricted families control only the miscalibration components they can approximate. Empirically, the debiased learned-critic objective achieves the best calibration among train-time calibration regularizers on four image and text benchmarks while maintaining competitive prediction accuracy, adds modest computational overhead, and remains complementary to post-hoc temperature scaling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.