The Predictable Core of Peer Review
Abstract
Peer reviewers often assign different scores to the same paper. Across 25,579 ICLR papers from 2020–2025, the Pearson correlation between one reviewer's score and the mean of the remaining reviewers ranges from 0.435 to 0.505. We show that collective ratings contain patterns that a small text model can learn. Our model combines word and character TF-IDF features with ridge regression over identity-filtered scientific text. On all 1,028 public NAIDv2 test papers, it reaches Spearman correlation 0.522, surpassing the dedicated NAIPv2-8B scorer's 0.439 and reducing calibrated mean-squared error by 7.03%. The comparison uses the same papers, outcomes, and calibration protocol, with broad manuscript text for the sparse model and NAIPv2's native title-and-abstract interface. Annual forward evaluations, presentation controls, and fixed-budget reading experiments trace the advantage to information beyond the opening text. Reading more of the manuscript reduces NAIPv2's added benefit for mean-score prediction by 34.8%. Human judgments supply further information after the manuscript model. Sparse learning thus reveals a predictable core in collective ratings while individual reviews contribute additional signal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.