acceptodds
Under review as a conference paper at ICLR 2027

TrajFix : Correcting Failed Trajectories to Scale Verifiable Software-Engineering Data

Abstract

Training an agent on real software-engineering (SWE) tasks requires trajectories that pass the tests. However, agents struggle to produce such a trajectory on hard tasks. Failed trajectories are normally thrown away, so most of the compute spent is wasted. Worse, the discarded trajectories come from the hard tasks and provide particularly valuable training signals. To make use of these failed trajectories, we present TrajFix, a multi-agent framework that repairs failed trajectories. It keeps the correct prefix and repairs only the first step that a diagnosis agent identifies as wrong, letting the agent continue the rollout until it submits a solution. Because the correction relies on oracle information that could leak the answer, a hallucination check identifies reasoning jumps and sends those corrections back for regeneration. Under a matched rollout budget per task, TrajFix improves the resolve rate by 44.0 points over resampling while reducing the cost of producing a trajectory usable for training by 47.6%. On large-scale GitHub tasks, we achieve a recovery rate of 53.60% and turn 19171 failed trajectories into training data. Compared with training on the original successful trajectories alone, adding recovered trajectories consistently improves SWE performance. For Qwen3-8B, SWE-Bench Verified improves by 5.25 points, SWE-Bench Multilingual by 6.16 points, and SWE-Bench Pro by 5.75 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.