Mitigating the Preference Trap of LLM User Simulators in Persuasion Dialogue RL
Abstract
LLM agents, despite strong performance on logical-reasoning tasks, still underperform in persuasion, a non-collaborative social interaction that demands adaptive strategies across a wide range of people over long-horizon conversations. Multi-turn reinforcement learning with LLM user simulators offers a promising paradigm to endow LLM agents with persuasion capabilities through open-ended interaction. However, we reveal that this paradigm suffers from a subtle training dynamic, which we term the preference trap: as training progresses, agents can be misled into abandoning broadly effective strategies in favor of brittle ones that exploit the training simulator's narrow preferences, earning high training reward but failing to generalize to unseen counterparts. Such exploits cannot be foreseen before training, because they only appear as the policy evolves. We therefore propose SimPatch, which detects exploited preferences during training and patches them through targeted instructions added to the simulator's prompt. SimPatch first identifies suspicious high-reward rollouts that share similar interaction patterns, then diagnoses the simulator-specific preference that makes the underlying strategy unusually effective. It then converts this preference into targeted instructions and representative examples, which are added to the simulator's prompt for subsequent rollouts. Without modifying the simulator weights, agent architecture, or RL algorithm, SimPatch provides a lightweight way to prevent repeated exploitation of simulator-specific preferences. On three held-out LLM persuadees, SimPatch raises GRPO's reward from 0.257 to 0.493 (Qwen3-4B) and from 0.283 to 0.391 (Qwen3-8B), outperforming all baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.