SEWISE: Signed Error-Weighted Informed Subgroup Estimation
Abstract
Many machine learning classifiers can achieve good average accuracy but sometimes fail systematically on subgroups of their inputs. Subgroup discovery methods (SDMs) surface such underperforming groups by using knowledge about which samples were misclassified. As common SDMs use a binary indicator to represent an error, when a model leaves one subgroup under-detected and its complement over-detected, the two error types cancel. For example, on CheXpert, where a pleural effusion classifier’s errors track patient sex, the two groups are far apart in signed error rate but nearly indistinguishable once direction is dropped. Additionally, the discovered errors are only part of a subgroup and therefore SDMs might miss the correctly predicted inputs. We present SEWISE, which represents each mistake as a signed error score, which is positive for false negatives and negative for false positives and is scaled by the model’s confidence, and diffuses it over a nearest-neighbour graph constructed from the classifier’s features to recover additional subgroup examples from their failing neighbours. Across ISIC-2017, CheXpert, CelebA and Waterbirds, SEWISE outperforms SPOTLIGHT, DOMINO and GEORGE on the majority of tasks. We further show that error-guided discovery may not be suitable when the classifier already has a good performance, as there is little error signal to utilise.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.