DynaEval: Large Language Models Struggle to Track Evolving Task States
Abstract
Large language model evaluation has largely focused on static, single-turn tasks. Real-world interactions, however, often involve multi-turn conversations in which task requirements evolve. We evaluate model performance in this setting and find that accuracy declines across dynamic multi-turn conversations. We attribute this decline to limitations in models’ ability to reconstruct the current task state from information dispersed across the conversation history. To evaluate this ability, we introduce DynaEval, a benchmark that simulates dynamic, long-horizon task environments. DynaEval comprises 543 eleven-turn sequences across mathematics, code, and databases, and provides two aligned versions of each turn: one presents an incremental update to the task conditions, while the other provides a complete description of the conditions valid after that update. Both versions share the same target question and reference answer. DynaEval comprises 543 eleven-turn sequences across mathematics, code, and databases. Under the full-history protocol with incremental updates, all ten models exhibit lower final-turn than first-turn accuracy, with the mean falling from 67.9% to 32.9%. Further diagnostic experiments attribute this cross-turn accuracy decline primarily to models’ difficulty in recovering the currently valid task conditions from information dispersed across prior interactions, revealing a bottleneck in multi-turn conversations and long-horizon tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.