acceptodds
Under review as a conference paper at ICLR 2027

Training Under an Unknown Future: Long-Horizon Task Agents That Leave Room for What Comes Next

Abstract

Agents are increasingly asked to implement successive requirements over long horizons, and a change focused solely on the current requirement may make later ones infeasible or substantially more expensive to implement. Existing training does not address this failure: rewards are assigned to the current task before downstream consequences become observable. On SlopCodeBench, which evaluates three to eight sequential requirements while requiring all earlier tests to keep passing, none of the fifteen evaluated frontier agents solves a complete problem. We propose a training method built on consequence chains mined from real development histories. Each chain links an early structural decision to later requirements that expose its downstream effects. We train on three-requirement chains, a practical length for exposing these effects. The objective first requires whole-chain test correctness, then refines the reward with a model-based judge's assessment of downstream impact, paying the model for writing code future steps can build on. On ChainSWE, our method improves the complete-chain success rate on the agent's own accumulated patches by at least relative to the corresponding base models, across scales from B to B. Our B model achieves a sequential bug-resolution rate of on ChainSWE, compared with for Claude Opus 4.6 under the same harness. On SlopCodeBench, checkpoint-level improvements extend beyond the three-step training horizon: on sequences of up to eight steps, the trained B model improves core solve rate by percentage points at steps four through eight, compared with points over the first three.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.