Towards Calibration for General Crammer-Singer Losses in Multi-Class Classification with Abstention
Abstract
The Crammer-Singer (CS) loss has been a standard multi-class extension of base binary surrogate losses. Despite its solid empirical performance, classical theory of loss functions proves that the CS loss is not classification-calibrated to the 0-1 loss, that is, the minimization of the surrogate risk of the CS loss does not necessarily lead to the minimization of the 0-1 risk. The pitfall of the CS loss is caused by low-confidence samples with their maximum class probability less than . This observation leads to the following question: is the CS loss calibrated when there are no low-confidence input samples? To this end, we consider the CS loss under multi-class classification *with abstention*, and in particular, the predictor-rejector framework, where a learner trains a class predictor as well as a rejector to abstain from making predictions for low-confidence samples. Under this setup, we show that general CS losses can be calibrated, verifying our hypothesis that the CS loss can be safely used even theoretically if all input samples are sufficiently confident. Theoretically, our calibration proof is interesting on its own right because it is applicable to broader losses composed of binary ones; the CS loss under classification with abstention has been considered previously, but the calibration property has been known only for polyhedral losses. Experiments on different datasets empirically validate the effectiveness of our calibrated general CS losses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.