acceptodds
Under review as a conference paper at ICLR 2027

ReplayWorld: Evolving Agentic World Model for Agent Training

Abstract

LLM-based environment simulators increasingly synthesize trajectories for agent training, reducing reliance on costly real environments. However, prompted and fine-tuned simulators generalize poorly across environments, and their environment knowledge stays fixed after deployment. We introduce ReplayWorld, an agentic world model for eight environments in web navigation, CoWork and tool use. It searches for information, executes code and consults a repertoire of simulator skills that encode how environments respond to actions. Replay-driven skill evolution writes recurring discrepancies between predicted and recorded observations into these skills, so environment knowledge evolves without weight updates. Semantic balanced sampling (SBS) promotes replay diversity, learnability-prioritized filtering (LPF) focuses selection on learnable traces, and multi-criteria cycle arbitration (MCCA) selects the best repertoire through independent validation. A hierarchical skill router (HiSkill) routes relevant skills during trajectory synthesis. On ReplayBench, our benchmark of 2,080 real environment transitions scored on six criteria, ReplayWorld outperforms its prompted backbone by 4.2 points. The agent trained on ReplayWorld trajectories gains 12.0 points on average across eight agent benchmarks. Code is available at https://anonymous.4open.science/r/replayworld-anon-2027-3ABB/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.