XYZ-SWE: Revisiting the Long Horizon Agent Training by Modeling-System Co-Design
Abstract
Long-horizon agent training, including software engineering, is advancing rapidly, but the recipe and infrastructure needed for a strong reinforcement-learning baseline remain difficult to reproduce. This makes gains from new algorithms or data hard to distinguish from improvements to an under-tuned baseline. We present XYZ-SWE, a fully asynchronous RL framework that establishes such a baseline through model-system co-design, using overfitting on a small task pool as the testbed. This testbed exposes failures that early reward gains conceal and guides the integration of established components: gated reward, importance-weight control, aligned training and inference engines, and consistent state across policy updates. No single component is new, yet together they form the missing foundation. On raw R2E-Gym data, without new algorithms or data work, trains Qwen3.5-4B to 60.5% pass@1 on SWE-bench Verified. This is comparable to concurrent FrogNano (61.5%), which relies on online task synthesis, after about one-eighth as many training rollouts. Building on this foundation, we further filter and refine R2E-Gym, SWE-rebench V2, and Scale-SWE. With this curated subset, the 4B model improves from 47.4%/25.0% to 62.3%/38.0% on SWE-bench Verified/Pro, and the recipe also improves Qwen3.5-9B from 56.2%/30.8% to 63.8%/40.6% on Verified/Pro. Controlled comparisons show that both curation and test strengthening improve held-out performance, although stronger tests reduce apparent training-task mastery. provides a documented starting point for studying better algorithms, data and environments for asynchronous agent RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.