acceptodds
Under review as a conference paper at ICLR 2027

DynaEval: Large Language Models Struggle to Track Evolving Task States

Abstract

Large language model evaluation has largely focused on static, single-turn tasks. Real-world interactions, however, often involve multi-turn conversations in which task requirements evolve. We evaluate model performance in this setting and find that accuracy declines across dynamic multi-turn conversations. We attribute this decline to limitations in models’ ability to reconstruct the current task state from information dispersed across the conversation history. To evaluate this ability, we introduce DynaEval, a benchmark that simulates dynamic, long-horizon task environments. DynaEval comprises 543 eleven-turn sequences across mathematics, code, and databases, and provides two aligned versions of each turn: one presents an incremental update to the task conditions, while the other provides a complete description of the conditions valid after that update. Both versions share the same target question and reference answer. DynaEval comprises 543 eleven-turn sequences across mathematics, code, and databases. Under the full-history protocol with incremental updates, all ten models exhibit lower final-turn than first-turn accuracy, with the mean falling from 67.9% to 32.9%. Further diagnostic experiments attribute this cross-turn accuracy decline primarily to models’ difficulty in recovering the currently valid task conditions from information dispersed across prior interactions, revealing a bottleneck in multi-turn conversations and long-horizon tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.