Quote or Invent? Fabricated Evidence in LLM-Written Peer Reviews
Abstract
Estimates indicate that language models now draft a visible share of peer reviews. Reviews can back their comments with quotations from the paper, i.e., words in quotation marks. The other reviewers and the authors trust a quotation in a review to provide the paper’s own words. When quotations turn out to be reworded or invented, that trust is lost, and with it also the trust in the quality of peer review. We measure how often this happens. Fourteen models from Anthropic, OpenAI, Google, DeepSeek, Moonshot and Meta review 100 accepted ICLR 2026 papers under a prompt that requires a verbatim quotation for every strength and weakness. In one condition the model receives the paper. In the other, the same paper with numbered sentences, and every quotation must cite the sentence it comes from. A deterministic checker locates every quotation in the text the model read, and three LLM judges from vendors other than the reviewing model’s grade every quotation the checker cannot confirm. 30% of the 2,759 usable reviews contain at least one quotation that is not the paper’s wording, but the models are far apart. With the paper alone, the rate runs from 2 and 3% for Gemini 3.1 Pro and GPT-6 Astra to 70% for gpt-oss-120b, and the share of reviews with a quotation whose meaning differs from the paper from 0 to 16 %. Numbering the sentences and asking for the number lowers the rate by 3.7 points. The quote-back check of a deployed review assistant would pass 73% of the invented quotations. The fuzzy matcher of earlier work accepts every changed number, negation and changed digit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.