acceptodds
Under review as a conference paper at ICLR 2027

THE TURN PARTITION IS A FREE PARAMETER IN MULTI-TURN LLM EVALUATION

Abstract

Large language models often score lower when a complete request is split across several conversation turns. Measuring this gap requires choosing how to turn the request into a conversation. We study how that choice affects the conclusions of a multi-turn evaluation. Using four tasks from the sharded benchmark of Laban et al. (2026), we evaluate five open-weight models across 24 conversation designs. These designs vary the number of turns, requirement order, the amount of information provided initially, and whether an earlier requirement is repeated, while preserving the final requirements and evaluation checker. We observe that conversation design substantially changes measured performance: the best and worst designs differ by a median of 0.123 on a 0-1 scale. Gaps between models can be small relative to their variation across designs, and the gains from four mitigation methods vary across designs, sometimes changing sign. When the same content is given in a single message, performance still varies across designs, indicating that design sensitivity is not exclusive to interactive conversations. These results show that conclusions about multi-turn performance and mitigation effectiveness are conditional on the conversation design used to evaluate them. We recommend evaluating across multiple explicitly specified designs and reporting variation alongside average performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.