acceptodds
Under review as a conference paper at ICLR 2027

InertiaBench: When LLMs Acknowledge Corrections but Still Give Old Answers

Abstract

Large language models (LLMs) have demonstrated impressive capabilities in multi-turn conversations, yet how they handle user corrections remains poorly understood. To investigate this, we perform experiments where a user revises the premises, goal, or constraints after the model answers and later asks the full revised question. We observe an intriguing phenomenon: models acknowledge a revised premise, goal, or constraint but still give the old answer, a failure we term thinking inertia, which accuracy-based benchmarks conflate with other errors. In this paper, we introduce InertiaBench, a diagnostic benchmark that examines whether models acknowledge a correction, apply it in later answers, and recover when reminded. Concretely, we pair problems from extensive datasets with LLM-generated corrections and standalone revised questions, cross-check both reference answers with distinct LLMs, and manually modify a stratified sample. Finally, we gather 5,996 instances spanning 4 domains and 3 correction types, and responses are classified as the old answer, the correct revised answer, or another error. To evaluate the InertiaBench performance of popular LLMs, we conduct comprehensive experiments on 17 models, illustrating that each model gives the old answer to 18–21% of revised questions, usually after acknowledging the correction, and that 43–77% of these failures persist after a reminder, more often for reasoning models (68%) than standard LLMs (54%). Our code and data are available at https://anonymous.4open.science/r/InertiaBench-4165.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.