acceptodds
Under review as a conference paper at ICLR 2027

Credit Follows Use: Used-Grouped Policy Optimization for Long-Context Memory Agents

Abstract

Long-context memory agents have emerged as a promising paradigm for retaining information across document streams and multi-session interactions. However, optimizing these agents poses a fundamental credit-assignment challenge: the consequences of a memory operation may remain unobserved until its content is retrieved and grounded in a distant response. Existing approaches typically either propagate a trajectory-level reward to all operations or derive process-level supervision from gold evidence and branched rollouts. We argue that both paradigms overlook an observable signal intrinsic to external memory systems—each materialized memory item exhibits a traceable lifecycle from writing, through retrieval, to grounded use in downstream generation. Inspired by this observation, we propose Use-Grouped Policy Optimization (UGPO), which records Write–Retrieve–Use traces and organizes policy updates around recurring item-use events rather than complete memory states. By modeling this lifecycle, UGPO routes task, use, and cost credit to responsible operations and propagates delayed outcomes to their originating writes. We further characterize state-group collapse, where free-form memory updates progressively eliminate state-level comparison groups, while item-use groups remain comparable across diverged memory states. Extensive experiments on long-document and multi-session benchmarks demonstrate that UGPO consistently improves task performance and attribution quality while reducing memory pollution and serving cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.