Architectural Priors for Ordinal Regression via Local Heads and Ordered Features
Abstract
Ordinal prediction problems—age estimation, disease grading, historical image dating—involve labels with intrinsic order, yet deep ordinal methods act at the loss or the output-side posterior, leaving the head’s mapping from features to logits agnostic to label order. We argue that the head itself is the third and right locus of inductive bias, and develop a theoretical framework around it. We introduce a family of label-axis structured heads in which each feature dimension affects only a local neighborhood of ordered labels, and prove that such heads operate through two complementary mechanisms: an ordered classifier mapping that localizes predictions on the label axis, and an implicit bias that drives the backbone toward ordered feature geometry through label-frequency filtering of the gradient field. Across 23 dataset–backbone cells spanning age estimation, histology and diabetic-retinopathy grading, and historical image dating (K=4 to 70), and eight ordinal baselines, our heads win or tie 177–182 of 184 Holm-corrected pairwise comparisons on MAE (128–145 of 161 on the ranked probability score) and are the only methods within 3% median gap of the per-cell best in every label-axis regime, while every baseline fails in some regime by 22–233% at worst. Mechanism-level analysis confirms the two mechanisms are dissociable: the Poisson-Unimodal baseline realizes only the second—it stays within 3–10% at large K—the best of any baseline in that regime—but collapses at small K and transfers a weaker representation under head swapping.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.