CoverSim: Coverage-Guided User Simulation for Task-Oriented Dialogue
Abstract
Task-oriented dialogue systems are increasingly evaluated with LLM-simulated users, enabling scalable testing across diverse user goals and profiles. Yet current approaches often reward human-likeness without verifying whether the simulator actually follows the user condition it is meant to enact. As a result, a conversation can look realistic while failing to test the intended behavior of the system. In this work, we ask what makes a user simulator good for testing task-oriented dialogue systems. We formalize simulation-based evaluation as testing, where each goal–profile pair is a test input, but it counts toward coverage only if the simulator faithfully executes it. This turns simulator improvement into a coverage-optimization problem: reduce inconclusive executions so more goal–profile conditions become testable. We then optimize user simulators directly for this objective using evolutionary prompt optimization, producing simulators that are better at testing the system, not merely imitating the user. We instantiate our automated testing framework using the data and user profiles from SimulatorArena and show that our simulator optimization strategy improves profile-following accuracy by up to 9.5% and more than doubles validated test coverage. We then show that standard TOD evaluation hides both agent failures and simulator failures, yielding a misleading picture of agent performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.