Sparsifying Label Smoothing to Better Calibrate Classifiers
Abstract
Label smoothing is widely used to reduce neural network overconfidence. Existing variants, whether uniform or adaptive, differ in how much smoothing mass they assign and how they weight it, but all spread that mass densely across the label space, so semantically plausible alternatives and unrelated labels alike receive probability mass. We argue that effective label smoothing depends not only on the smoothing strength but also on the support over which smoothing mass is assigned: smoothing should be sparse and similarity-aware. We instantiate this principle with Sparse Similarity-Aware Label Smoothing (SSA-LS), a simple and practical method that uses latent-space class neighborhoods to redistribute probability mass only toward plausible alternatives. Across multiple datasets and architectures, SSA-LS consistently improves confidence calibration while preserving, and sometimes improving, accuracy. Beyond SSA-LS, similarity-aware sparsity acts as a general wrapper, improving the calibration of existing adaptive label smoothing methods without hurting accuracy. We further show that strong representations do not by themselves guarantee well-calibrated predictions, and through sparsity ablations and neighborhood corruption studies, that the gains arise specifically from structured, similarity-aware sparsity, establishing it as a critical ingredient for label smoothing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.