: Scaling Turing-Style Evaluation to Optimize Human Simulation in Dialogue
Abstract
Realistic user simulation is essential for developing and evaluating conversational AI agents, yet measuring and optimizing simulation fidelity remains challenging. Human Turing tests assess whether simulated behavior appears human, but are costly to use for iterative optimization, motivating scalable evaluation with Machine-Turing judges. We introduce , a human-grounded co-evolutionary framework for improving user simulators and Machine-Turing judges through prompt optimization, without updating model parameters. Machine-Turing judges provide rewards and diagnostic feedback to improve simulators, while human Turing judgments and explanations guide judge refinement. The refined judge then provides feedback for further simulator improvement. In customer-service dialogue, simulator optimization improves Turing win rates under both held-out LLM judges and human annotators. Incorporating human explanations improves held-out judge accuracy, and optimizing against the refined judge further improves simulator human-perceived human-likeliness. These results demonstrate the effectiveness of \turing in jointly improving user simulators and their evaluators through human-grounded co-evolution. The dataset would be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.