acceptodds
Under review as a conference paper at ICLR 2027

LLM-REVal: Can We Trust LLM Reviewers Yet?

Abstract

The rapid advancement of large language models (LLMs) has encouraged their integration into academic workflows, potentially reshaping how research is conducted and reviewed. In this study, we examine how the joint use of LLMs in research generation and peer review may affect scholarly fairness, focusing on the risks of using LLMs as reviewers through simulation. Our simulation incorporates a research agent that generates and revises papers, alongside a review agent that evaluates the submissions. Based on the simulation results, we conduct human annotations and identify notable misalignment between LLM-based reviews and human judgments: (1) LLM reviewers assign more favorable scores to LLM-authored papers than to matched human-authored papers; (2) some limitation-oriented human-authored papers remain below the acceptance threshold even after multiple revisions. Additional analyses further characterize these disparities through paper-level linguistic profiles, the effect of LLM polishing, and limitation-oriented topics. These results highlight risks and equity concerns for human authors and scholarly evaluation if LLMs are deployed in peer review without adequate caution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.