acceptodds
Under review as a conference paper at ICLR 2027

ICM-bench: A Benchmark for Evaluating the In-context Modification Ability of Large Language Models

Abstract

Large language model (LLM) agents continually accumulate conversation histories, yet the contextual state relevant to the current situation is not identical to the full history: some previously valid information may become obsolete, while other information should remain unchanged. We formalize the ability to selectively transform such contextual information during inference, without parameter updates, as In-Context Modification (ICM). Existing benchmarks on unlearning and long-term memory primarily evaluate forgetting, retrieval, or conflict resolution, but do not systematically test whether models can update the current contextual state while preserving unaffected information. To fill this gap, we introduce ICM-Bench, a controlled benchmark for evaluating ICM in multi-turn conversations. It contains over 2.8K multi-turn conversations and 26K queries, covering deletion and update across single-attribute, multi-hop, and entity-level settings, together with explicit retain and unmodified evaluations. All evaluated knowledge is purely contextual, isolating ICM from parametric knowledge. Experiments across different LLMs reveal a striking gap between using and modifying context: unmodified success ranges from 94.50% to 100.00%, whereas modified success ranges from only 0.25% to 35.75%. Existing inference-time methods can improve modification success, but often at the cost of retained or unmodified information. These results show that reliable ICM requires selective contextual state transformation rather than indiscriminate forgetting, and establish ICM-Bench as a benchmark for developing LLM agents that can maintain an accurate current state across evolving conversations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.