acceptodds
Under review as a conference paper at ICLR 2027

GeMS-SC: Learning a Cross-Attentive State-Skill Policy for Runtime Graph-Memory Control in Long-Horizon Agents

Abstract

Long-horizon LLM agents with graph-structured memory must decide at every execution step whether memory is needed, which graph-native behavior to invoke, and how returned structures should condition later actions. Existing approaches primarily optimize memory representation, retrieval, or generic entry-level operations, leaving this runtime control problem under-specified. We formulate graph-memory invocation as a state-conditioned skill-selection policy and introduce GeMS-SC. Its controller, the Cross-Attentive State-Skill Encoder (CASE), jointly contextualizes execution-state fragments, including the current subtask, previous observation, and recent trajectory, with candidate Graph-Native Memory Skills. A lightweight interaction encoder followed by Skill-to-State Cross-Attention constructs state-conditioned skill representations; a shared scoring head then selects one skill or Continue Regular Execution. The selected skill injects an operation protocol specifying the memory objective, graph-tool scope, argument cues, result handoff, and stopping condition, while the frozen base agent retains control over concrete tool calls and task actions. CASE is learned in two stages: weak state-skill alignment from future memory-tool behavior, followed by advantage-weighted trajectory optimization using task success, skill-behavior alignment, memory-operation quality, and execution efficiency. Controlled comparisons with state-aware LLM skill routing and direct graph-tool prompting isolate the effects of learned selection and intermediate skill guidance. Across Qwen3-8B-Instruct and Llama-3.2-3B-Instruct on LoCoMo, LongMemEval, ALFWorld, and MemoryAgentBench, GeMS-SC is competitive with the strongest structured-memory baseline on LoCoMo and yields consistent gains on cross-session and unseen interaction metrics, with its largest improvement on dynamic memory maintenance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.