acceptodds
Under review as a conference paper at ICLR 2027

HISTORY IS NOT STATE: CAUSAL DIAGNOSTICS FOR SUFFICIENT-STATE FAILURES IN LANGUAGE AGENTS

Abstract

Language agents are increasingly deployed in interactions where user goals, ev-idence, and environment states evolve over time. Final accuracy alone does not reveal where failures enter the history-to-state-to-decision chain. We pro-pose MarkovLens, a causal diagnostic framework that uses task-sufffcient states as intervention targets: canonical, non-leaking representations under which the target decision should be independent of the surface history, given the current query and decision context. We introduce four diagnostics: the sufffcient-state gap (SSG), Markov invariance violation (MIV), state update error (SUE), and obsolete-state attraction (OSA). On a 96-state balanced User Constraint Update grid, qwen2.5:7b improves from 0.833 raw-history accuracy to 1.000 gold-state accuracy (SSG = +0.167, 95% CI [0.130, 0.206]), while mistral:7b ex-hibits the opposite compact-state pattern (0.917 to 0.833, SSG = −0.083, 95% CI [−0.151, −0.023]). This opposite sign is not a contradiction: interface-equalized controls show that explicit-state rendering is itself part of the mechanism. Cross-task results extend the diagnostic beyond user constraints: on a 48-group Code Requirement Drift slice, qwen2.5:7b again shows a positive SSG (0.755 to 1.000, SSG = +0.245, 95% CI [0.135, 0.354]), while mistral:7b remains di-rectionally negative but statistically unstable. Tool/Environment Mutation shows that compact states can hurt while dialogue-style states can recover performance, and OSA calibration supports strong stale-state attraction for qwen2.5:7b but not equally for mistral:7b. MarkovLens is therefore a diagnostic microscope rather than a leaderboard or universal repair: it localizes whether models fail to infer the current state, fail to use an explicit state interface, or remain attracted to stale states.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.