GROVE: Benchmarking and Training LLM Agents to Persuade
Abstract
LLM agents increasingly settle technical disputes with one another before LLM judges that cannot verify every step, so the outcome turns on persuasion: winning the judge's preference over a competing argument. Recent work trains models to persuade, but a judged win rate cannot tell whether a trained agent also follows the task's procedure or supports the correct decision. We introduce GROVE-Bench, three environments in which agents rebut a code review, debate a factual claim with retrieved evidence, or challenge a solution to a graduate-level question. Each fixes what every participant sees, pits trained agents against frozen models from 2B to 397B parameters in complete round robins, and records judged preference separately from execution and dataset labels. We also propose GROVE (Group-Relative Optimization of Verbal Exchanges), which rewards each argument relative to others sampled for the same input without querying the judge, and instantiate it with programmatic rewards (GROVE-P). On items used during training or model selection, the SFT-DPO-GROVE-P pipeline raises field win rates over the base model by 46.9% in Code Review (9B, under a judge that labeled no training data) and 70.2% in Evidence Debate (4B), matching or beating every frozen model except the 397B one; in code, imitating winning arguments supplies most of the gain. On held-out reasoning questions, where GROVE ran only as probes of at most 25 updates, no method improves on the base model. In evidence debates, higher win rates do not imply better execution: the highest-rated trained debater leaves the required first search to the engine in almost every debate, and its wins do not track which side of the claim is true. Persuasion training should therefore be evaluated on execution and agreement with dataset labels as well as on judged preference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.