acceptodds
Under review as a conference paper at ICLR 2027

EchoRL: Policy-Evolving Environment-State Caches for Agentic Reinforcement Learning

Abstract

Agentic reinforcement learning (RL) trains policies that act inside real software environments, such as checking out repositories, installing dependencies, running builds, and invoking tools. Group-based learning algorithms, such as GRPO, sample multiple rollouts per task, causing the same deterministic environment operations to be repeated across rollouts and thereby incurring a considerable fraction of the end-to-end rollout latency. Caching environment operations is a natural solution, but policy evolution across training steps raises two challenges: (i) cached execution results may become invalid under updated context, producing incorrect trajectories and biased policy gradients; and (ii) the cache must continually adapt to retain the environment states most likely to be revisited by future policies. To this end, we present EchoRL, a runtime layer that reuses environment work. We make two contributions. First, we reduce environment latency by casting reuse as caching a deterministic transition function keyed by logical state and tool version, and show that under an explicit determinism assumption cached and uncached rollouts follow the same distribution, introducing no additional bias into the underlying policy-gradient estimator. Second, we introduce advantage-guided cache retention to anticipate environment-state reuse under evolving policies, improving cache hit rates in subsequent training iterations. Across all evaluated workloads, EchoRL significantly reduces per-rollout environment latency by – compared with No Cache while preserving trajectory correctness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.