TurnFlow: Evaluating LLMs Beyond Task Completion in Dynamic Multi-Turn Interactions
Abstract
Large language models (LLMs) are increasingly used for complex collaborative tasks that require users to add constraints and revise plans over multiple turns. Model responses, in turn, shape subsequent user behavior and task progression. These dependencies call for evaluation of both task outcomes and the interaction process. However, predefined user inputs in static evaluations cannot reflect reactions to model responses, while an emphasis on final success in dynamic evaluations can obscure differences in completion efficiency, turn-level progress, and the value of proactive guidance. To address this gap, we introduce TurnFlow, a benchmark for evaluating LLMs on complex multi-turn tasks. Derived from 250,100 real-world conversation sessions, TurnFlow comprises 400 tasks covering single-goal and multi-goal settings, short and long conversations, and scenarios requiring backtracking and revision. From the same user population, we derive 16 communication style configurations that characterize how users formulate initial requests and follow-up utterances. To evaluate models on these tasks under different communication styles, we develop a dynamic multi-turn evaluation framework comprising a path planner, a user simulator, a trajectory recorder, and an evaluator. The planner releases stage-specific task information and behavioral requirements as the conversation progresses. The simulator generates user reactions based on the conversation history, user-side task information, and communication style, allowing subsequent interactions to adapt to model behavior. The recorder tracks model responses, task progress, and user feedback. Using this evidence, the evaluator assesses model performance through four trajectory-based metrics: Task Completion Rate (TCR), Task Completion Efficiency (TCE), Single-Step Task Contribution (SSTC), and Beyond-Goal Value (BGV). Together, these metrics measure task completion, completion efficiency, turn-level contributions, and value beyond prior user expectations. Our results show that models can complete tasks without consistently advancing them effectively throughout the interaction, with limitations in proactively guiding users and providing value beyond their existing needs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.