TraceGen: Synthesizing Difficulty-Controllable Bugs for Training and Evaluating Code Repair Agents
Abstract
Training and evaluating code repair agents both demand large volumes of quality bug data, but existing synthesis methods depend on known patch locations and produce bugs without difficulty labels, leaving training data flat and evaluation unable to distinguish agents at different capability levels. We present TraceGen, a framework that mines structured causal paths, termed DefectChains, from repair traces and uses them to synthesize bugs at independently discovered code locations without access to the ground-truth patch. The causal depth of each DefectChain yields three difficulty tiers that serve a dual purpose: ordering trajectories for curriculum fine-tuning and producing calibrated evaluation gradients. On SWE-bench Verified seeds, TraceGen validates 194 of 318 attempts at a 61.0% rate, while baselines deprived of the original patch location suffer 63-92% validation drops. An agent trained via curriculum SFT on TraceGen trajectories achieves 24.7% on SWE-bench Verified at the 7B scale, surpassing SWE-Dev at 23.4%, R2E-Gym at 19.0%, and Lingma SWE-GPT at 18.2%; curriculum ordering alone contributes 3.3 percentage points over random-order SFT on identical data. All six evaluated frontier models exhibit monotonic resolve-rate increases across the three tiers, confirming the stability of the difficulty gradient.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.