DynaCAST: Intervention-Grounded State for Long-Horizon Language Agents
Abstract
Long-horizon language agents need compact persistent state that preserves distinctions in past experience relevant to future decisions. Existing memory objectives emphasize retention or prediction without specifying which state differences should matter when goals, world conditions, or plans change. Learning such a criterion also requires comparing open-ended language actions across histories and keeping component-specific targets consistent with a global recursive metric. We introduce \method, which grounds persistent goal, world, and plan registers in controlled futures generated by executable goal, physical-world, and route interventions in resettable environments. We make these futures comparable through canonical action classes and frozen argument-binding distributions, and derive global and component Bellman targets from shared transport couplings and a common maximizing control, ensuring exact target compatibility while establishing contraction and covered-value preservation for the exact operator. Experiments across eight benchmarks with nine matched baseline/control comparisons across four model families—Qwen2.5-7B, Llama-3.1-8B, Gemma-2-9B-IT, and Mistral-7B-Instruct—isolate –-point macro-averaged task-score gains from matched finite-horizon Bellman bootstrapping over direct Monte Carlo metrics in both state architectures, with intervention and register ablations identifying the contributions of executable controls and modular state, and exact checks on 24 enumerated delivery SCMs validating operator implementation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.