BranchAid: Pricing Search Against Trajectory Recovery in Repository Agents
Abstract
On 300 SWE-bench Verified issues, a repository agent solves more tasks by selectively repairing its live trajectory than by spending the same marginal credit on another search branch: GPT-4o moves from 30.2% to 31.7% solved while its solved-task median cost falls from $4.73 to $3.92 and superseded build/test invocations fall from 11.0 to 8.5; Claude 3.5 Sonnet moves from 33.1% to 34.8%, $4.41 to $3.64, and 10.4 to 7.8. The recovery exchange makes this allocation measurable through paired end-to-end runs with identical backbones, tools, 40-step and $15 caps, and event-level credit accounting; Extra-Search spends each released credit on exploration, while CostRule shares the recovery action menu and observable state. BranchAid implements the recovery policy with a heterogeneous graph of repository execution, calibrated operational hazards, action-conditioned terminal value, cost-aware ranking, and abstention. Across 900 matched issue–seed pairs per backbone, issue-clustered solve intervals favor BranchAid over both controls, and every paired cost and superseded-CI interval favors BranchAid. The gain concentrates on pre-defined hard issues—3.3 points for GPT-4o and 3.1 for Claude—while easy and mid-band outcomes remain statistically aligned, locating the measured advantage at long or CI-intensive trajectories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.