MESA: Scoring the Volume and Stability of Comparative Claims
Abstract
The rapid pace of AI/ML research produces a steady stream of methods claiming state-of-the-art performance, each backed by a leading benchmark score. However, reported scores offer practitioners limited guidance on adoption, as they are obtained under one configuration of the many ad-hoc choices an experiment requires. This leaves two questions unresolved: (1) the amount of the plausible configuration space in which a method outperforms others, and (2) whether that lead survives a small change to the configuration. Existing evaluation tactics fall short of answering these questions, as they either evaluate methods under fixed configurations or analyze each method in isolation. We introduce Meta-Experiment Stability Analysis (MESA), which answers both. MESA evaluates a set of candidate methods across plausible experimental configurations to determine the regions where each method is in the top- after accounting for the noise between repeated runs. MESA scores each method's top- region by (1) its volume, the plausibility-weighted fraction of the space it covers, and (2) its stability, the typical distance from a point in the region to its boundary. The same meta-experiment can then be reused to evaluate how well benchmark datasets separate the methods. We demonstrate MESA through two case studies from different areas of AI/ML. In both, no method leads on a majority of the configuration space on any dataset. These results suggest that a leading benchmark score is weak evidence that a method will lead under a practitioner's own choices, and that MESA's volume and stability scores give evaluation a more informative target.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.