How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
Abstract
As large language models increasingly participate in scientific evaluation, we investigate how rhetorical rewriting can reward-hack AI reviewers when reported scientific content is largely preserved. We construct 4,200 full-paper manuscripts from 120 anonymized ICLR 2026 submissions. Two LLM rewriters modify six rhetorical dimensions in opposing directions under scientific-content preservation constraints, and five LLM reviewers assess the manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, followed by scope framing, with this hierarchy persisting across human-rating ranges. More elaborate rewriting does not reliably yield larger gains: joint-rewrite gains depend strongly on the rewriter, reviewer guidance does not consistently outperform an unguided second pass, and recursive gains vary across configurations. Strict review lowers mean overall assessment by 1.36 points, but positive-negative rhetorical contrasts persist. These findings reveal selective, configuration-dependent vulnerabilities in AI scientific review and motivate evaluation systems robust to rhetorical variation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.