acceptodds
Under review as a conference paper at ICLR 2027

RepetitionBench: A Conversational Benchmark for Multi-Turn Agent Repetition

Abstract

Repetition is a pervasive failure mode in conversational Large Language Model (LLM) agents, manifesting as re-asked questions, redundant confirmations, and unnecessary restatements. Such behavior disrupts the natural flow of the dialogue and erodes user trust. Yet repetition remains under-measured, and existing benchmarks do not stress-test this behavior in realistic, conversational settings. To address this gap, we introduce RepetitionBench, a benchmark to quantify and classify different repetition modes in phone-based healthcare conversations between humans and AI agents. RepetitionBench comprises three components: (i) a fine-grained taxonomy spanning semantic, structural, and surface-level dimensions of repetition; (ii) 13 verifiers (6 deterministic and 7 subjective) to judge the repetition of an LLM agent; and (iii) an adversarial dataset of 6,892 samples drawn from healthcare conversations averaging over 20 turns. We publish a total of 7,516 labeled samples, including 18,372 human annotations of repetition types and 8,736 verifier outputs. Across 13 frontier models, surface-level repetition is rare, but semantic repetition remains common, especially in longer conversations. RepetitionBench is the first work to enable a systematic study of repetition in multi-turn conversational LLM agents, providing a basis for scalable evaluation prior to real-world deployment, while the verifiers in this work can be used to evaluate and analyze different modes of repetition.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.