Web2LongTraj: Scaling Long-Horizon Agent Trajectories from Natural Web Corpora for Mid-Training
Abstract
Long-horizon agent trajectories are expensive to obtain because most existing pipelines synthesize them in manually engineered executable environments. We present Web2LongTraj, a corpus-first pipeline that mines natural web corpora for real multi-step problem-solving records and compiles them into source-grounded agent trajectories, with actions and outcomes constrained by the source and observations audited for machine-output form. Screening 317 million web records, Web2LongTraj yields 1.15 million audited trajectories. To scale these short episodes into long-horizon supervision, we introduce two operators that grow them along the two axes on which real sessions become long. AnchorStretch extends dependency distances within a trajectory: it first identifies an early premise and a later step that depends on it, then inserts topically similar audited episodes between them. ArtifactStitch composes multi-task sessions by walking a graph that connects trajectories through shared artifacts and compatible working contexts. By recombining the 1.90B-token audited pool, the two operators generate 10.6B long-context tokens, a 5.6-fold expansion without building new task environments. Mid-training Olmo-3-7B and Olmo-3-32B on Web2LongTraj with a standard next-token objective yields consistent gains over equal-budget baselines on tool calling, multi-step interaction, long-horizon memory, and coding benchmarks under both Instruct and Think post-training recipes, while preserving general capability. On Olmo-3-7B, the full Web2LongTraj recipe improves all eight benchmarks by an average of 4.7 points under Instruct SFT and 9.9 points under Think SFT; its five-benchmark general-capability mean also increases by 0.4 and 3.3 points, respectively. These results establish corpus-first trajectory construction as a practical complement to environment-first synthesis for training long-horizon agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.