Repairing Failed Coding-Agent Trajectories for Cost-Efficient Training Data Expansion
Abstract
Coding-agent trajectories provide supervision for model training and evidence for failure analysis, yet collecting them through repeated rollouts is costly and often produces unsuccessful runs. Existing curation pipelines largely retain solved trajectories or regenerate complete runs, leaving action-level reuse of failed trajectories insufficiently explored. We introduce , a framework for augmenting coding-agent trajectory datasets by recomposing failed runs. localizes anomalous actions, replaces them with compatible spans from successful trajectories, and verifies the resulting repository states. For the first research question (RQ1), the strongest checker reaches 93.72 anomaly F1, while a compact 3B variant retains 87.77 F1 and powers the remaining pipeline. In RQ2, 18.23% of failed trajectories pass build and executable functional tests, and 11.78% yield verified alternative implementations, outperforming the evaluated local-replacement baselines. In RQ3, three-seed training on a 16,000-record source-verified, state-aligned corpus reduces held-out action negative log-likelihood by 22.05% relative to the donor-context baseline. A separate resource-normalized evaluation finds that the marginal cost of acquiring an accepted output is 3.9–210 lower than complete rollout with the evaluated agents. These results show that trajectory recomposition recovers useful training supervision from failed coding-agent runs without generating complete replacement rollouts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.