Faithful by Construction: Hierarchy-Aware Decision Rules for Better Mistakes
Abstract
Hierarchical classification considers settings where label classes are the leaves in a tree , called a taxonomy, and leverages that taxonomy to reduce mistake severity: a misclassified image should land on a similar class and not an unrelated one. Existing methods build the taxonomy into the training recipe, through a loss, soft targets, or feature regularization, but lack a principled account of precisely how these changes affect the decision rule. We ask what it means for a classifier's decision rule to incorporate taxonomic distances exactly. We define prototype rules, classifiers that decide by measuring distances to class prototypes in feature space, and -faithfulness, the property that squared distances between class prototypes match taxonomic distances exactly. We characterize which feature layouts and classifier heads achieve -faithfulness, and propose a simple frozen-head training objective that realizes -faithfulness by construction. An empirical evaluation shows that -faithful models offer competitive accuracy and the lowest mistake severity among several strong baselines. Through a neural collapse analysis of a simple mean-squared-error training objective, we identify two roadblocks that make it difficult to train a -faithful model without freezing the classifier head: we must anticipate and account for the effective backbone regularization strength in advance of training, and the distance-based decision rule is distorted by class-dependent offsets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.