acceptodds
Under review as a conference paper at ICLR 2027

Multiple Task-State Descriptions Interfere with In-Context Computation

Abstract

Adding a second description of a task state lowers language-model accuracy on a downstream question about it, even when the two agree and even when one is an exact copy of the other. When one description updates the other, the model fails on most items where it can separately resolve the current state and compute from it, once it must do both together. We call this a composition penalty. Neutralizing one token’s representation in the obsolete description restores most of the lost accuracy, showing that the obsolete state remains causally active. Duplication and stale-state costs recur in verbatim shell-tool output, and on an independently written benchmark the penalty persists on items whose state the model resolves correctly, including for Claude Opus 5.5, which solves every controlled chain. In nineteen novels, deleting the sentences that support an obsolete state improves accuracy more than deleting as much other text. Reasoning closes some controlled gaps but leaves a gap on long histories. Compiling the current state into a single description before computing nearly eliminates the penalty in the controlled setting. Its end-to-end benefit depends on compilation fidelity: on long histories, inaccurate compilations can outweigh the benefit and reduce overall accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.