Generalized Example Difficulty Metrics and Their Applications for Tabular Classification
Abstract
Quantifying example difficulty in ML is crucial for addressing prevalent data quality issues. Our goal is to design the most effective example difficulty metric for each data-centric application. To this end, we investigate questions about the design space itself: Which design axes matter most? Which values are the "best" for each axis, given all else equal? Existing surveys, which are catalogs of published metrics, cannot answer these questions rigorously. Furthermore, existing research under-represents the pervasive domain of tabular data, and none rigorously contrasts the architecture families of deep learning and gradient-boosted decision trees (GBDTs). We remedy these limitations by defining a broad taxonomy that generalizes the design knobs of existing metrics, yielding 96 metrics in a five-dimensional design space, spanning DL and GBDTs. We then study their behavior across 33 tabular classification tasks and their performance across five data-centric applications. We show the following: 1) Most published metrics correlate together despite a vast design space. 2) Model probing method matters more than other axes, such as whether to compute influence or direct train- or test-time measurements. 3) Final-model and snapshot-mean measurements are redundant, and train-time and test-time measurements are nearly redundant. 4) Training and measuring GBDTs is more effective for application tasks on tabular data than the DL equivalent on average; loss remains strong for many tasks. We conclude by identifying the best metric, from our experiments, for five data-centric application tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.