Better Reviews, Busier Area Chairs: Why a Little Human Judgment Matters in AI-Assisted Review
Abstract
Growing submission volumes have intensified interest in AI-assisted peer review, but its benefits depend on the expert effort required to turn reviews into reliable decisions. We study the downstream workload of Area Chairs (ACs), whose capacity to integrate reports and determine what requires further assessment can become a bottleneck. We develop a Bayesian framework linking AI-assisted reports to the AC's downstream workload, defined as the minimum additional assessment cost required to estimate paper quality at a fixed reliability target. We establish a paradox: AI-assistance can improve both individual report accuracy and the panel's aggregate assessment while increasing AC workload, even when the AC chooses assessments adaptively and optimally. We further show that a small amount of independent reviewer judgment may be sufficient to eliminate this increase and substantially reduce AC workload. Experiments using expert-annotated real reviews from PeerReview Bench are consistent with these theoretical findings. Our results highlight the value of preserving independent human judgment when designing AI-assisted review processes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.