Most Highlighted Comparisons Cannot Be Checked From the Printed Page
Abstract
Results tables in machine learning often signal the strongest method by boldfacing a value, but such claims can be independently assessed only when the paper reports sufficient uncertainty information. We study how often this is possible using random samples totaling 2,895 papers accepted at NeurIPS and ICML in 2024-2025. An automated pipeline identifies highlighted results and their strongest competitors, interprets table structure and metric direction, and determines whether each comparison contains enough information for statistical testing. In the 2025 sample, only 16.9% of highlighted comparisons are testable from the reported uncertainty and run counts. Among these testable comparisons, 47.6% of highlighted differences are not statistically distinguishable from zero using Welch's -test at the stated run counts, and this fraction increases after correcting for multiple comparisons within tables. The 2024 sample exhibits the same overall pattern, and sensitivity analyses across alternative statistical and reporting assumptions do not change the qualitative conclusion. Validation against manually checked reference sets supports the reliability of the extraction pipeline, and the resulting data and evaluation artifacts are released with the paper. Our findings suggest that uncertainty and run counts should be reported at the level of the comparisons used to support empirical claims.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.