FAIL2TASK: Reproducing Coding-Agent Failure Patterns through Executable Task Synthesis
Abstract
Failed coding-agent trajectories provide evidence of weaknesses encountered in practice. However, when the original repositories and execution environments are unavailable, these failures cannot be directly replayed for evaluation or agent improvement. We study failure-to-task reconstruction: turning such trajecto- ries into new executable tasks rather than recovering the original tasks. We introduce FAIL2TASK, a closed-loop framework whose goal is to reproduce the source failure pattern when the same agent attempts the constructed task. Evidence-supported failure patterns guide repository selection and task construc- tion, producing tasks with runnable environments, tests, and reference solutions. Execution-based validation and semantic comparison of source and replay trajec- tories guide iterative refinement. We evaluate FAIL2TASK on 96 source failures from SWE-bench Verified and Terminal-Bench 2.0, producing tasks accepted by the pipeline for all 96 cases. Rerunning the agents on these tasks yields trajec- tories with 80.18% failure-pattern alignment with the original failure trajectories. Harness self-evolution on the constructed tasks achieves source-task pass rates of 27.1–31.3% on SWE-bench, compared with 14.6–20.8% for the strongest base- line. On Terminal-Bench, the corresponding rates are 16.7–20.8%, compared with 4.2–12.5% for the baselines. These results show that tasks constructed from fail- ure trajectories can help agents solve previously failed tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.