acceptodds
Under review as a conference paper at ICLR 2027

Dep-Search: Learning Dependency-Aware Reasoning Traces with Persistent Memory

Abstract

The evolution of agentic systems has increasingly relied on reinforcement learning (RL) and persistent memory mechanisms to enable autonomous decision-making and long-term knowledge retention. Recent advances in agentic RL have demonstrated the potential for agents to learn complex behaviors through interaction with their environments, while persistent memory systems have shown promise in maintaining and reusing accumulated knowledge across multiple tasks. However, existing approaches often struggle with efficiently searching through vast action spaces and managing dependencies between learned behaviors and stored memories. In this work, we introduce Dep-Search, a dependency-aware search framework that integrates agentic RL with persistent memory to enable efficient exploration and knowledge reuse. Dep-Search employs a structured search strategy that explicitly models dependencies between actions and learned experiences, allowing agents to navigate complex state spaces more effectively. The framework leverages persistent memory to maintain and retrieve relevant experiences, while RL components learn to optimize search policies based on dependency relationships. By organizing search around dependency-aware structures, Dep-Search enables agents to accumulate and reuse knowledge across episodes, rather than treating each search instance independently. Experiments on multi-task learning and sequential decision-making benchmarks demonstrate that Dep-Search achieves improved sample efficiency and task performance compared to existing agentic RL methods. Further analyses reveal that the dependency-aware search mechanism leads to more interpretable agent behaviors and effective knowledge transfer through persistent memory. These results highlight the importance of combining structured search strategies with memory mechanisms in advancing the capabilities of agentic RL systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.