TRACE-SWE: Trajectory Refinement with Action-oriented Execution Evidence for Code Agents
Abstract
Effective supervised fine-tuning (SFT) of code agents benefits from trajectories that connect intermediate actions with execution feedback. Real-user interaction histories from code agents such as Claude Code provide these records. However, they interleave useful actions with redundant exploration and failed attempts, limiting their value as direct training supervision. We introduce TRACE-SWE (Trajectory Refinement with Action-oriented Execution Evidence for Code Agents), a framework that refines real-user trajectories into action-oriented supervision using recorded execution evidence. For successful trajectories assessed as completing user requests, TRACE-SWE uses recorded outcome evidence to retain relevant actions and their supporting context while removing unnecessary steps. Inspired by cloze-style learning, TRACE-SWE then masks individual actions or action sequences and reconstructs them from preceding and subsequent context. For failed trajectories with unresolved user requests, TRACE-SWE uses recorded failure evidence to construct corrective supervision. TRACE-SWE improves performance on SWE-bench Verified, Pro, and Multilingual across four open-weight models. On LING-3.0-FLASH, it achieves an average gain of 6.70 percentage points across six SWE-bench configurations under THINK inference, with improvements also extending to Terminal-Bench. TRACE-SWE also reduces average interaction rounds and non-cached tokens per solved task, indicating improved efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.