acceptodds
Under review as a conference paper at ICLR 2027

SCoPR: Scaling State-Evolution Horizons for Long-Horizon Agents

Abstract

Coding LLMs are evolving from single-turn code generators into software-engineering agents that operate over extended horizons in real repositories. As agent-authored PR trajectories such as OpenClaw grow rapidly, a central question emerges: can software-engineering trajectories generated by coding agents, after rejection sampling, be used to train stronger long-horizon agents, and can scaling trajectory horizon yield measurable capability gains? We introduce SCoPR to generate candidate rollouts with a chain-of-PR pipeline and select high-quality chains by measuring how each intermediate state in the functional evolution aligns with the golden patch. To enrich the state-evolution process, we introduce two auxiliary pipeline improvements: world-model execution feedback, which mitigates rollout interruption caused by unavailable or unstable code execution, and diff-preserving handoff, which prevents state information loss during workspace transitions. We find that rejection-sampled agent-authored trajectories provide effective supervision for state-evolution-oriented training data. Moreover, under controlled data budgets, increasing the effective horizon yields additional capability gains. These results suggest that agent-authored code data, when combined with human-authored software data, can simultaneously expand the available training resource and improve downstream performance. More importantly, the trajectory horizon extends the scaling axis of long-horizon agents beyond raw token quantity toward cumulative state evolution and persistent goal pursuit, providing a more effective path for unlocking the intrinsic potential of long-horizon agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.