acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Mean Accuracy Comparisons in Pedestrian Attribute Recognition

Abstract

Pedestrian attribute recognition (PAR) methods are commonly compared using mean accuracy (mA), the macro-average of per-attribute balanced accuracy. Our paper-and-code audit finds heterogeneous label-imbalance interventions paired with a small set of numerical inference thresholds. Because these interventions can change score meaning, native mA comparisons can involve different operating policies. We develop a decision-coordinate framework that derives objective-implied coordinates for identifiable loss and sampling components and estimates trained predictors' score-space coordinates through validation-set isotonic calibration. The framework complements native mA with objective and realized-policy normalization toward the common balanced-accuracy decision target. We evaluate controlled methods with ResNet-50 and ViT-B/16 backbones and published PAR methods, holding each predictor fixed across evaluation views and using disjoint test data. On RAP, the gap between inverse-prior reweighting and unweighted binary cross-entropy decreases from to percentage points with ResNet-50 and from to with ViT-B/16 under realized-policy normalization. Comparisons among published methods using the same backbone also exhibit reversals in mean ordering, while some gaps persist or widen. These results show that the magnitude and direction of method advantages can depend on operating policy, supporting joint reporting of native and normalized mA.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.