acceptodds
Under review as a conference paper at ICLR 2027

When Does Decodable Task Information Constitute Agent State? Dissociating Information, Currentness, and Behavioral Use in Live Coding Agents

Abstract

Probing is widely used to study task-relevant information in a model's internal representations, including information about the state of an ongoing task. For an acting agent, however, treating such information as state raises two further questions. Does it track the task as it changes, and does it shape the actions the agent takes? We study these questions in coding agents working on real software repositories, recording model activations and using task variables whose values can be determined directly from the environment. To separate current state from recency, we create conflicts in which an observation from the environment provides the current value, while stale evidence presented later asserts a different one. Our results empirically dissociate decodable information, authoritative currentness, and behavioral use. In our primary setting, where the prompt explicitly identifies the authoritative source, the current task value remains strongly decodable and continues to guide the agent's natural action. Yet a readout trained on ordinary cases fails at the frozen decision threshold when stale evidence is presented last. When current and stale values vary independently during fitting, however, the resulting readout transfers to held-out conflict data in every eligible case we test. In a separate construction, authority must be inferred from provenance rather than explicit labels, and we vary presentation order. The readout of the current value remains strong regardless of whether current or stale evidence appears first. We complement these measurements with causal interventions that separate changes in an internal score, in the decision process, in the emitted action, and in the final action. Effects depend on where the intervention enters the computation, and an action changed at one decision point may still be restored later in the trajectory. Together, these results show that evidence for agent state extends beyond representation to how task information enters decisions and persists as the task evolves.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.