acceptodds
Under review as a conference paper at ICLR 2027

StateSwap: Separating Propagation from Decision Effects in Persistent-State LLM Agents

Abstract

Evaluations of persistent-state agents can conflate three questions: whether an intervention was executed, whether its consequences propagated through retrieval to rendered evidence, and whether the downstream decision responded. StateSwap verifies intervention execution separately and reports, on paired latent worlds, a state-intervention profile : allocation disagreement, rendered-evidence disagreement, observed decision disagreement, and the paired correctness contrast. Its Intervention–Propagation–Decision Separation principle states that rendered-evidence disagreement requires allocation disagreement, whereas complete upstream disagreement does not identify a decision-side effect. We instantiate this protocol with real, externally served readers. Within each pair, the task, plans, repository realisation and paired gold target are fixed; only the stored allocation policy differs, inducing a versus split of a fixed retrieval budget, and the reader is masked from state, arm and gold. In a pre-specified held-out block of 2000 paired synthetic worlds, allocation and rendered evidence changed in every pair while (95% interval -0.003 to +0.023), compatible with a small positive, near-zero or slightly negative effect. An annotation-assisted calibration on 1000 HotpotQA distractor-validation questions under two reader families shows why this distinction matters: two one-paragraph replacements with identical upstream change rates have sharply different decision-side consequences. The mean contrast for replacing an unannotated paragraph relative to baseline met a pre-specified pp equivalence criterion under both readers, whereas replacing an annotated supporting paragraph reduced accuracy by 34.5 pp (DeepSeek) and 32.2 pp (Qwen). A separate fresh-sample validation on 3200 new, R9-disjoint synthetic worlds tested an imposed stored policy. For the pinned Qwen primary, the reliability-dependent decision-side interaction met the prespecified conjunctive sign-reversal criterion; the prespecified DeepSeek secondary did not reach the reversal. The design was fixed after the exploratory R9 outcome was known, so it is not a preregistration or an independent replication, and S35's held-out identity is unchanged. StateSwap therefore treats propagation and decision response as distinct estimands; it neither measures model-internal state nor establishes a general decision benefit from persistent state.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.