acceptodds
Under review as a conference paper at ICLR 2027

EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent

Abstract

The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across stateful workspaces. While Agentic Reinforcement Learning (Agentic RL) provides a promising path to optimize these agents, its scaling is heavily bottlenecked by the severe scarcity of interactive training environments. Existing synthetic environments are strictly limited to tool-calling endpoints, rendering them insufficient for accommodating the end-to-end real-world demands of claw-like agents. To bridge this gap, we introduce ClawForge, an automated framework for synthesizing executable environments and scalable training data. Specifically, ClawForge employs an environment synthesis engine to build sandbox-isolated workspaces, alongside a topology-aware data generation engine to produce coherent task trajectories. Overall, we synthesize 139 interactive environments comprising approximately 20K complex tasks for Agentic RL training. Experiments on Qwen3/3.5 models (8B–32B) show that ClawForge yields gains of up to +11.9% on Claw-style benchmarks and +8.0% on general tool-use benchmarks, with concurrent reductions in inference token cost. The result confirms that synthesized executable environments

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.