StateCross: A Benchmark for Memory and State Tracking in Long-Horizon Agents
Abstract
Long-horizon tasks require agents to follow instructions that may change across sessions. An instruction may be withdrawn, take effect later, or apply only to one part of a task. Agents must therefore remember information and decide which requirements govern their next action. Existing memory evaluations often focus on recall or answer consistency. These measures do not by themselves show whether retained instructions lead to correct actions. Evaluations centered on memory can also overlook workspace files, which preserve instructions and earlier results after the conversation is cleared. We introduce StateCross, a benchmark for evaluating whether retained information supports actions that follow current requirements. Controlled coding tasks combine changing requirements with context resets while workspace files and explicit memory persist. Agents select from predefined programs, whose executed effects are checked against the requirements in force. We compare agents with and without explicit memory using decision accuracy, requirement-specific errors, and consistency across consecutive actions. The framework also distinguishes evidence gaps at storage, retrieval, and model input, supporting separate analysis of information availability and action correctness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.