ARD-Review: Generating Accurate, Reasonable, and Diverse AI Peer Review
Abstract
Current AI review systems tend to generate generic, unsuited, nitpicky, and invalid (GUNI) feedback, and this failure mode risks homogenizing not only reviews but, through defensive writing, the papers themselves. We propose ARD-Review, a split-role pipeline that combines a high-recall harsh critic, an adversarial merger that aggressively filters GUNI content, and a retrieval-based calibration step that anchors scores to human-reviewed reference papers. We also introduce a new evaluation suite and benchmark to quantify the GUNI tendency in AI-generated reviews. On a leakage-controlled dataset of ICLR 2026 pre-rebuttal reviews, our method outperforms prior systems in scoring accuracy with the least sycophancy bias, is the only method significantly better than the human one-vs-one baseline, and wins 80% pairwise validity comparisons while keeping cross-paper overlap close to the direct-review level. SycoPaperEval: 46 poor papers (average score ) reviewed by three models known by the community to be scyphantic or contrarian toward papers due to their training. We further show that the same calibration mechanism can recalibrate individual human scores and can substantially reduce the spherical nature of backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.