VerMem: Verifier-Guided Command-Level Credit Assignment for Unified Memory Management
Abstract
Long-horizon language model agents must preserve useful knowledge, manage bounded context, and recover earlier task evidence. In existing memory policies that broadcast trajectory-level advantages, the shared signal does not directly assess the semantic quality of individual commands. We introduce VerMem, a unified memory policy with verifier-guided, command-level credit assignment. The local branch evaluates executable memory transitions and normalizes their semantic scores across commands of the same operation type within an update. The global branch normalizes composite feedback across candidate trajectories for the same task. Combining these advantages with direct constraint costs provides differentiated credit while retaining the task-level objective. The policy optimizes complete commands with equal command-level weight and coordinates versioned long-term memory, active context, and same-task history. Supervised initialization and staged reinforcement learning develop these capabilities. Both semantic verifiers provide training feedback and are disabled during evaluation. Experiments on question answering and interactive tasks show improved overall task performance, stronger results from the combined feedback branches, and better performance under constrained online-token budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.