ToolTurns: Benchmarking LLM Agents in Multi-Turn Dialogue
Abstract
Tool-using language models are increasingly deployed in conversational settings, yet most large-scale tool-use benchmarks focus on isolated, single-turn requests. We introduce ToolTurns, a benchmark for evaluating language agents end-to-end in multi-turn dialogue with tools. It contains 2,620 problems spanning 1,827 tools and 11,004 API endpoints across 47 categories. Each problem begins with an intentionally underspecified request and unfolds into three user queries, yielding 7,860 independently evaluated subtasks that require agents to seek clarification, maintain conversational context, select tools, and respond to follow-up requests. We additionally introduce an open-source user simulator that reveals withheld context through dialogue and generates context-dependent follow-ups, together with a checker–fixer pipeline for improving simulated user turns. Across nine language models, the best agent resolves 69.6% of individual subtasks but completes only 39.8% of full three-query dialogues, and every evaluated model performs worse on follow-up queries than on the initial request. We further find that dialogue-quality rubric scores and task completion can rank models very differently, highlighting the importance of evaluating whether users' underlying requests are actually resolved. Finally, paired turn-level evaluation shows that our revision pipeline reduces automatically detected dialogue defects by 41.2%, while blind human evaluation prefers revised turns in 63.3% of cases. These results expose substantial room for improvement in reliable, multi-turn tool use and provide a scalable testbed for measuring progress.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.