acceptodds
Under review as a conference paper at ICLR 2027

ECHO-Graph: Scaling Long Horizon RL with Dependency-Aware Credit Propagation

Abstract

Long-horizon agentic reinforcement learning requires not only extending the accessible interaction history, but also assigning effective learning signals across repeatedly reconstructed contexts. ECHO couples context management with credit assignment by selecting historical memories during reconstruction and reinforcing the turns retained in the final reconstruction trace. However, final-trace masking on successful trajectories can leave early prerequisite steps unreinforced as reconstruction depth grows, biasing learning toward later stages. Fully masking unsuccessful trajectories removes corrective feedback, risking rapid entropy collapse and reduced exploration. We propose ECHO-Graph, which propagates outcome credit backward through a DAG of turn dependencies induced by context reconstruction, reinforcing early prerequisite steps and repeatedly reused evidence in successful trajectories. For unsuccessful trajectories, where provenance does not reliably support fine-grained credit assignment, we apply attenuated dense penalties to provide corrective feedback and help sustain exploration. The method requires no additional rollouts, critics, or model evaluations. Experiments on BrowseComp-Plus and SWE-Bench Pro show gains over ECHO, especially on difficult instances, with stronger context-horizon scaling on the former.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.