acceptodds
Under review as a conference paper at ICLR 2027

MemCredit: Reinforcement Learning for Memory-Based Search Agents via Retained-State Credit Assignment

Abstract

Memory-based search agents repeatedly retrieve external evidence, compress it into a compact memory, and condition subsequent search and reasoning on this retained state. However, existing reinforcement learning objectives are dominated by terminal answer rewards and provide little supervision on whether a search turn actually leaves the agent with a more useful memory. As a result, retrieving relevant evidence does not necessarily imply retaining the information needed by later decisions, creating an acquisition–retention gap. We introduce MemCredit, a reinforcement learning framework that jointly redesigns the intermediate reward and its credit assignment for memory-based search agents. First, MemCredit defines a retained-state reward that evaluates each search turn through the change in gold-answer answerability of the agent’s compact memory. Specifically, a frozen reference model scores consecutive memory states, and their potential difference provides a dense signal for whether the current search-memory transition improves the information actually available to future decisions. Second, we develop a memory-aware credit assignment mechanism that converts this retained-state reward into a bounded, group-relative local advantage and combines it additively with the standard trajectory-level outcome advantage. The resulting credit is assigned to the complete search-memory conversation including planning, querying, and memory updating rather than only to the query or memory-writing tokens, while final-answer learning remains governed by the terminal outcome reward. This design also supplies discriminative local supervision when sibling trajectories receive identical terminal rewards. Experiments show that MemCredit outperforms all the baselines over seven datasets. Additional analyses investigate the acquisition–retention gap and examine how retained-state rewards and conversation-level credit assignment contribute to the performance gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.