acceptodds
Under review as a conference paper at ICLR 2027

Chain of Pull Requests: Unlocking Long-Horizon Agency Data-Efficiently

Abstract

While Large Language Models (LLMs) excel at short-term tasks, scaling them to long-horizon agentic workflows remains challenging for the scarcity of data that captures authentic long-dependency structures and cross-state evolutionary dynamics. Existing synthesis methods either limit tasks to single-feature scenarios or incur prohibitive human annotation costs, failing to provide scalable, high-quality supervision. We address this by reconceptualizing data synthesis through the lens of real-world software evolution. Our key insight: PR sequences naturally embody the very supervision signals needed for long-horizon learning. They decompose objectives into units, maintain iterative coherence, and encode bug-fix refinements. Building on this, we propose Chain of Pull Requests (CoPR), which mines structured supervision from PR sequences through three interlocking mechanisms: (1) progressive task decomposition via continuous commits, (2) long-term consistency enforcement through unified functional objectives, and (3) verifiable refinement from authentic bug-fix trajectories. CoPR preserves causal dependencies and iterative refinements, naturally aligning with full-cycle task modeling. The resulting trajectories are substantial—averaging 85k tokens and 116 tool calls—yet remarkably data-efficient: fine-tuning GLM-4.6 on 239 CoPR samples yields broad benchmark improvements, notably achieving a 47% relative gain on Toolathlon. Further analysis confirms the effective internalization of long-horizon behaviors and unveils training and inference scaling for extended horizons. Our work establishes CoPR as a scalable paradigm beyond limitations of single-feature synthesis, proving that modeling real-world evolutionary trajectories offers a principled path to unlock the intrinsic long-horizon potential of agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.