Scaling Traces Diversity with Runtime-Decoupled World Model for Mid-training
Abstract
Frontier models increasingly incorporate agentic experience earlier—from post-training into mid-training—to build fundamentals in agentic behaviors and environment interaction. This stage teaches broad patterns of exploration, action, and recovery before specialized agentic SFT and RL, creating demand for both high-quality trajectories and diverse, distinct tasks. Existing execution-based pipelines provide high-quality traces but rely on curated, runnable environments and verifiable outcomes. In agentic coding, this concentrates supervision on resolved issues and merged pull requests, leaving open issues, development branches, and tasks with unusable images or non-reproducing gold solutions underrepresented. Our re-execution audit estimates that roughly one-third of SWE-rebench V2 tasks fall outside reliable execution-based supervision, while over 55M unresolved public issues offer an additional source of tasks without known solutions. Motivated by these gaps, we use Qwen-World-Model to simulate execution-dependent feedback while retaining exact repository state for deterministic actions. This enables trajectory collection without requiring every task to have a working image or known solution. We construct a corpus spanning 3M unique tasks and 100 B training tokens, increasing unique-task coverage by 4× at 4× the throughput of execution-based collection. In matched-budget SFT, replacing 25%, 50%, or 100% of executed traces with simulated traces yields comparable performance across four agent benchmarks, supporting their utility as training data. At matched token exposure, adding traces from new queries outperforms repeating the original collection by 4.6 pass@1 points on average, indicating that broader task coverage is more valuable than repeated rollouts at a fixed token budget. Further incorporating these diverse traces into mid-training improves pass@1 after the same SFT by +8.7, +6.2, +2.1, and +1.1 percentage points on SWE-bench Verified, Multilingual, Pro, and Terminal-Bench 2.0, respectively. These results suggest that world-model simulation can expand both the volume and breadth of agentic experience, strengthening the foundations learned during mid-training and improving downstream agent performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.