Toward Valid Agent Evaluation: Reconciling Task Fidelity and Human Plausibility in LLM-Based Interactive User Simulation
Abstract
LLM-based user simulators enable scalable evaluation of interactive agents, but task success is informative only when the simulated interaction itself provides valid evidence of agent capability. This can fail in two ways: the simulator may drift from the assigned task or it may exhibit behavior that is extreme relative to real users. These risks motivate our joint assessment of Task Fidelity (TF) and Human Plausibility (HP). We audit 30 user LLMs, finding that 34.1% of successful interactions contain HP violations and 45.9% of failures contain TF violations. An intuitive response is to prompt simulators with personas, which we find improves HP but often at the expense of TF. To address the risks, we propose a method that separates what the simulator needs to accomplish from how it pursues those goals. Across 30 user LLMs, this separation turns average TF loss into a gain (−6.4 → +0.6 pp), while further improving HP (+6.1 pp). These gains extend to agent evaluation, with reduced score sensitivity to user LLM choice, more consistent rankings, and strong agreement with external evaluations. Together, these findings show that task-faithful, human-plausible interactions provide a stronger basis for interpreting agent outcomes as evidence of capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.