Search-Mem: Learning to Search via Reinforcement Learning with Memory Agent
Abstract
Reinforcement learning (RL) has recently been used to train search agents that iteratively retrieve evidence for large language models (LLMs). However, two problems remain. First, a search agent can stop before gathering enough evidence to produce the correct answer. Once the search stops, the system has no way to recover from insufficient or incorrect evidence. Second, deeper search increases the amount of retrieved information carried across search steps, which can eventually exceed the available context and remove evidence needed for later decisions. We introduce Search-Mem, a memory-augmented multi-agent search system trained with RL. The system consists of five cooperating agents: a planner, compressor, retriever, generator, and memory agent. The planner controls the search by deciding what to query, which retrieved documents to keep, when to use the compressor, and when to stop. The compressor is trained jointly with the planner to condense retrieved evidence into short statements that can be carried into later search steps. Both agents are optimized using the improvement that the gathered context brings to the answer of a frozen generator, while the retriever and generator remain fixed. After RL training, the system constructs the memory agent by replaying multi-hop questions and storing how each question was decomposed without using reference answers. At inference, the trained search agents first gather evidence, and the generator produces several answers. If these answers are inconsistent or empty, the memory agent retrieves similar past decompositions and uses them to guide an additional search. The additional search gives the system another chance to find missing evidence, while the compressor keeps the planner's context short as the search continues. Across six QA benchmarks and three frozen generators, Search-Mem achieves the best average accuracy under both an open and a proprietary judge, with the best result in 34 of 36 settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.