RMPO: Replay-Measured Turn-Level Credit for LLM Agent Training
Abstract
Multi-turn agent training often relies on sparse episode rewards, making it hard to determine which actions contributed to the final outcome. Standard solutions either infer credit from trajectory statistics or rely on compute-heavy rollout sampling. In this paper, we show that credit can be measured more directly: by restoring the execution state in the environment, applying a localized change to the action, and evaluating the difference in outcome. We formulate this procedure as a common replay interface and propose *Replay-Measured Policy Optimization* (RMPO), a critic-free method that adds normalized replay measurements to the episode-level advantage. Depending on what the environment supports, RMPO implements three pragmatic credit operators: leave-one-out action deletion, prefix-progress scoring, and same-state action replacement. We evaluate RMPO in ALFWorld, WebShop, and Search-QA with Qwen3-4B and Qwen3-8B. It consistently outperforms GRPO and GiGPO. In ALFWorld, deletion measurements correlate strongly with symbolic-planner progress (), and RMPO exceeds four-continuation sampling at equal GPU-hours (90.5% versus 85.5%). The method requires a resettable replay environment. Code is available at https://anonymous.4open.science/r/rmpo-CD48.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.