What Matters in Dataset and Benchmark Papers? A Large-Scale Analysis
Abstract
To investigate the influence and importance of different factors in dataset and benchmark publications, we first systematically analyze 1,615 papers accepted to NeurIPS Datasets and Benchmarks Track between 2021 and 2025, combining OpenReview metadata, Semantic Scholar citations, GitHub repository statistics, and information extracted from paper PDFs. We then extend our analysis to the 3,951 papers concerning benchmarks and datasets published at NeurIPS, ICML, and ICLR. We consider numbers about the papers, such as the author counts and citations, as well as numbers in the papers, such as dataset sizes and the number of tasks. Papers with more authors generally receive more citations, and the number of stars and forks on GitHub is positively correlated with citation counts. Special recognition such as spotlight or oral strongly implies higher, even after controlling for paper's underlying quality using a regression discontinuity design. However, dataset size and the number of baselines show little correlation with citation counts. We hope these findings encourage a more rational use of such numbers when writing and reviewing dataset and benchmark papers, and we release our curated dataset to support further research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.