Large Language Models Learn to Persuade Multiple Receivers through Social Context
Abstract
Large language models (LLMs) enable open-ended persuasion, yet existing LLM-based approaches largely focus on isolated, short-horizon interactions, overlooking how target selection and social influence jointly shape collective outcomes. To address this gap, we formulate LLM-based multi-receiver persuasion as a natural-language sequential decision problem, in which an LLM persuader decides whom to address and what to say to heterogeneous receivers connected by a social graph. We instantiate this setting in three domains: academic rebuttal, political decision-making, and charitable campaigns. Our key insight is that evolving social context is central to both persuasion evaluation and policy learning. For evaluation, we find that incorporating social context generally improves correlation and reduces prediction error with human judgments across all three domains and accordingly develop Persuade-RM, which outperforms the task-specific Rebuttal-RM in academic rebuttal. For policy learning, we propose Graph-Contextual Group Policy Optimization (GCGPO), a critic-free algorithm that groups socially comparable steps and combines step-level advantages with episode-level relative advantages. Across all three domains, GCGPO achieves strong quality and stance scores while avoiding the degradation observed in several baselines. We further develop a coupled actor-reward generation mechanism that reduces training time per iteration by 44.7% on average while maintaining comparable evaluation performance. Together, these results establish social context as a key ingredient for effective evaluation and policy learning in LLM-based multi-receiver persuasion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.