When Can an AI Reviewer Replace a Human? Human-Equivalence Selective Review
Abstract
As AI systems spread through the research pipeline, peer review faces pressure to scale without losing human judgment; generating a plausible critique is different from reproducing how human reviewers decide. We study 454 ICLR submissions from 2024–2026 with 1,810 public human reviews and evaluations from four foundation models. The sharpest reliability gap occurs in the half-level region around the human recommendation boundary: moving outside this region improves AI–human agreement, although rare extreme scores are not monotonically safer. This observation yields Human-Equivalence Selective Review (HESR), a fixed assignment rule that lets one AI review occupy a review slot when its ordinal score lies at least half a level from the human-aligned boundary and otherwise routes the paper to a person. HESR requires no training, model-specific parameters, batch ranking, or multi-model inference. Across 99 deduplicated papers and 396 model–paper judgments, HESR retains 54.8% of reviews and raises agreement with individual human recommendations from 67.6% to 73.6% (paper-cluster bootstrap 95% CI for the gain, 2.4–9.8 percentage points). Historical decision accuracy increases from 73.5% to 87.1%; all four base models and all six evaluation blocks improve, with 109 accept-side and 108 reject-side assignments. Two frozen tests reproduce the effect on 40 unseen papers. On a quality-controlled cross-year test, the gain remains positive after adversarially removing one human review from every paper. These results show that a single AI reviewer can match the variability of human recommendations on an identifiable subset of papers while preserving human review for the least stable cases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.