acceptodds
Under review as a conference paper at ICLR 2027

Trace-MT: A Coverage-Guided Benchmark for Evaluating Complex Multi-Turn Dialogues

Abstract

Multi-turn conversations between real users and LLMs often exhibit complex and diverse behaviors such as interleaving, loops, and interruptions. Such behaviors make every response contingent on the conversational state accumulated over all prior interactions, yielding a combinatorial space of state dependencies that existing single-turn evaluations cannot capture. In this paper, we introduce Trace-MT—to the best of our knowledge, the first coverage-guided benchmark for evaluating complex user behaviors in multi-turn dialogues. Similar to a control-flow graph, we formalize a multi-turn dialogue as a dialogue-flow graph, where each node is a task (totally 6 types of tasks) and each edge denotes the transition taken in response to the source task. As such, we can leverage the concept of coverage, which has been widely used in control-flow testing, for evaluating complex multi-turn dialogues. To avoid exponential explosion as the number of turns increases, we define an N-gram on such dialogue-flow graphs as a sequence of N nodes. We set (equal to the number of task types) and, by enumerating all candidate traces and applying pruning algorithms, obtain a set of 472 specific traces. By selecting one topic to cover each trace, we build a benchmark of 472 coverage-guided tests, where the instructions for simulating user responses are dynamically generated from each trace. The benchmark results on nine open-weight models reveal interesting findings and demonstrate the usefulness of the proposed benchmark.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.