Uncovering Robustness Weaknesses in LLM Agents via Persona Evolution with Task-Integrity Constraints
Abstract
Interactive-agent evaluations commonly rely on standardized LLM-based user simulators that exhibit broadly cooperative behavior, potentially overlooking failures that arise under diverse but realistic user interactions. This can lead to overly optimistic estimates of agent robustness and provide limited failure signals for training agents that generalize across users. Existing approaches largely rely on predefined user behaviors, which may miss agent- or domain-specific weaknesses, while failure-oriented search can reduce task success by invalidating the underlying task rather than exposing genuine robustness failures. To fill this gap, we introduce Tiper, a task-integrity-preserving persona evolution method that automatically discovers agent-specific user policies exposing robustness weaknesses while keeping the underlying task valid and solvable. Tiper evolves persistent user personas through reflection-guided behavioral mutations and task-wise Pareto selection under a task-integrity constraint. Across four interactive-agent benchmarks, Tiper yields lower held-out avg@3 scores in all 12 GPT-5.4 settings, with observed reductions of 5.8-33.7 percentage points, and in all 28 agent–domain pairs spanning seven agents, with a median observed reduction of 22.9 points. Further analysis yields a behavior-to-failure taxonomy that connects discovered user behaviors to recurring agent failure modes. Together, these results show that Tiper enables more diagnostic robustness evaluation of interactive agents and provides informative failure analysis for future agent robustness improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.