acceptodds
Under review as a conference paper at ICLR 2027

Most Highlighted Comparisons Cannot Be Checked From the Printed Page

Abstract

Results tables in machine learning often signal the strongest method by boldfacing a value, but such claims can be independently assessed only when the paper reports sufficient uncertainty information. We study how often this is possible using random samples totaling 2,895 papers accepted at NeurIPS and ICML in 2024-2025. An automated pipeline identifies highlighted results and their strongest competitors, interprets table structure and metric direction, and determines whether each comparison contains enough information for statistical testing. In the 2025 sample, only 16.9% of highlighted comparisons are testable from the reported uncertainty and run counts. Among these testable comparisons, 47.6% of highlighted differences are not statistically distinguishable from zero using Welch's -test at the stated run counts, and this fraction increases after correcting for multiple comparisons within tables. The 2024 sample exhibits the same overall pattern, and sensitivity analyses across alternative statistical and reporting assumptions do not change the qualitative conclusion. Validation against manually checked reference sets supports the reliability of the extraction pipeline, and the resulting data and evaluation artifacts are released with the paper. Our findings suggest that uncertainty and run counts should be reported at the level of the comparisons used to support empirical claims.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.