Trace-MT: A Coverage-Guided Benchmark for Evaluating Complex Multi-Turn Dialogues
Abstract
Multi-turn conversations between real users and LLMs often exhibit complex and diverse behaviors such as interleaving, loops, and interruptions. Such behaviors make every response contingent on the conversational state accumulated over all prior interactions, yielding a combinatorial space of state dependencies that existing single-turn evaluations cannot capture. In this paper, we introduce Trace-MT—to the best of our knowledge, the first coverage-guided benchmark for evaluating complex user behaviors in multi-turn dialogues. Similar to a control-flow graph, we formalize a multi-turn dialogue as a dialogue-flow graph, where each node is a task (totally 6 types of tasks) and each edge denotes the transition taken in response to the source task. As such, we can leverage the concept of coverage, which has been widely used in control-flow testing, for evaluating complex multi-turn dialogues. To avoid exponential explosion as the number of turns increases, we define an N-gram on such dialogue-flow graphs as a sequence of N nodes. We set (equal to the number of task types) and, by enumerating all candidate traces and applying pruning algorithms, obtain a set of 472 specific traces. By selecting one topic to cover each trace, we build a benchmark of 472 coverage-guided tests, where the instructions for simulating user responses are dynamically generated from each trace. The benchmark results on nine open-weight models reveal interesting findings and demonstrate the usefulness of the proposed benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.