REPO: Replay-Enabled Policy Optimization for Terminal Agents via Trajectory-Reconstructed Environments
Abstract
Realistic, executable environments are scarce for terminal-agent post-training, while recorded trajectories are plentiful but remain fixed demonstrations. We introduce REPO (Replay-Enabled Policy Optimization), which reconstructs an executable workspace from each trajectory by deterministic replay and agentic completion, then adds per-turn container snapshots for counterfactual rollouts. The recovered workspace supports intent recovery, new single- and cross-workspace tasks, and multi-round sessions, each paired with an executable verifier. It expands each workspace along breadth (cross-project dependencies) and depth (persistent user follow-ups). We further introduce PAPO, which anchors sibling rollouts at verifier-diagnosed pivotal turns, and CBAO, which branches at high-entropy turns to obtain local credit signals. On public trajectories, REPO yields k task-sufficient environments. SFT warm-start followed by PAPO+CBAO improves Qwen3.5-27B by 11.9 points on Terminal-Bench 2.1 and by 18.3 points in EvoCode-Bench v2 MT@4.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.