acceptodds
Under review as a conference paper at ICLR 2027

ASSERT: Synthesizing Replay-Verified Trajectories with Access Isolation for Software Engineering Agents

Abstract

Reference patches specify verified fixes, but do not provide the reliable interaction traces needed to train software engineering agents. Rejection sampling discards failed attempts, while patch-conditioned traces can contain unsupported reasoning and redundant steps. Therefore, we introduce ASSERT, a framework that converts reference patches into executable training trajectories across programming languages. Its central design separates privileged plan recovery from the execution evidence used to justify supervised decisions. Dependency-Constrained Trajectory Repair (DCTR) repairs failed plans within dependency regions identified by failure feedback, while Replay-Certified Decision-Sufficient Slicing (RCDS) removes redundant interactions subject to evidence, replay, and final-test constraints. A patch-blind runner supplies observations, and an isolated rationale writer explains fixed actions using preceding evidence. With GPT-5.5 and GLM-5.2 teachers, ASSERT retains up to 2.93 times as many trajectories as the largest baseline pool and improves student Pass@1 by up to 5.84 percentage points over the strongest baseline in each setting on SWE-bench Verified and Multi-SWE-bench. To compare supervision quality at matched task coverage, we also train a student for each method using one method-specific trajectory per task in each setting's four-method training intersection. ASSERT maintains a Pass@1 advantage of up to 6.07 percentage points over the strongest baseline and has the lowest unsupported-claim rate in every setting, supporting improvements in both supervision yield and training utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.