When Labels Have Structure: Improving Image Classification with Hierarchy-Aware Cross-Entropy
Abstract
Standard cross-entropy is the default classification loss across virtually all of machine learning, yet it treats all misclassifications equally, ignoring the semantic distances that a class hierarchy encodes. We propose Hierarchy-Aware Cross-Entropy (HACE), a drop-in replacement for standard cross-entropy that incorporates a known class hierarchy directly into the loss. HACE combines two components: prediction aggregation, which propagates the model's probability mass upward through the class hierarchy so that parent nodes accumulate the confidence of their children; and ancestral label smoothing, which distributes the ground-truth signal along the path from the true class to the root. We evaluate HACE on CIFAR-100, FGVC Aircraft, and NABirds in two regimes: end-to-end training across six architectures spanning convolutional and attention-based designs, and linear probing on frozen DINOv2-Large features. In end-to-end training, HACE improves top-1 accuracy over standard cross-entropy in 15 out of 18 architecture–dataset pairs, with a mean gain of 5.40 percentage points. In linear probing, HACE outperforms all competing methods on all three datasets, with a mean improvement of 2.18 points over the next best baseline. Beyond leaf-level accuracy, HACE also makes better mistakes: its predictions are more accurate at coarser levels of the taxonomy. Evaluated with hierarchical metrics on NABirds, it achieves the best operating curves and the smallest distance to the ground truth in the taxonomy in 5 of 6 architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.