acceptodds
Under review as a conference paper at ICLR 2027

SWE-Trail: History-Guided Trajectory Collection for Coding Agent Training

Abstract

Coding agents solve repository issues through long sequences of exploration, editing, and validation. These interaction trajectories are increasingly used as post-training data, so how they are collected shapes what the trained model learns. A common approach is to sample multiple trajectories independently for each issue and then select or curate them. However, as more trajectories are sampled independently, later trajectories become increasingly redundant. In fact, earlier trajectories can guide where subsequent collection effort should be spent. A natural way to use this history is to expose it to the coding agent as advice, helping it decide what to do next. But training on such trajectories can make the model rely on advice that is unavailable at test time, creating a train-test mismatch that may reduce downstream performance. Instead, we argue that history should guide *where* collection effort is spent, while the coding agent decides *what* to generate by itself. Based on this principle, we introduce **SWE-Trail**, a trajectory collector that uses earlier trajectories to guide subsequent collection while keeping that history outside the coding agent's context. During a new rollout, SWE-Trail uses history to identify a state where additional collection effort may be useful. When that state is reached, the coding agent proposes candidate next actions from its ordinary context. SWE-Trail briefly executes candidates from the same checkpoint and uses the results to decide which route continues. We compare collection approaches by training student models on the trajectories they produce and measuring downstream performance. For a 124B student initialized from Ling-3.0-flash-base, SWE-Trail improves mean resolution over independent sampling by 2.33 and 1.60 percentage points on SWE-bench Multilingual and SWE-bench Pro, respectively. Without history, the same collection pipeline produces more resolved trajectories than independent sampling, but improves by only 0.11 points on Multilingual. These results show the value of using history to guide collection without exposing it to the coding agent. Our code is available at https://anonymous.4open.science/r/SWE-Trail-3D13/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.