plorer: Scaling Coding Agents through Reconciled Tree Search
Abstract
Test-time scaling is a promising way to improve LLM-based coding agents. However, long-horizon software repair produces extended trajectories whose partial progress must be preserved, compared, and reused; simply adding steps or branches can yield diminishing returns. We present plorer, an iterative tree-search framework built around a repeatable explore–reconcile–prune cycle. Within each iteration, path-level best-first selection extends promising trajectories; at the iteration boundary, compatible paths exchange evidence and the frontier is compressed to seed the next iteration. This bounded design lets the search use larger test-time budgets while repeatedly consolidating its most promising progress; end patch selection then compares candidates accumulated across the full search. On SWE-bench Verified with a gpt-5-mini backbone, increasing the test-time budget from 30 to 50 improves SWE-Xplorer from 54.6% to 61.4%, while the prior tree-search framework SWE-Search improves from 54.2% to 54.8%. At a matched budget of 50, SWE-Xplorer leads SWE-Search by 6.6 points while reducing average cost and runtime. It also achieves a higher resolved rate than strong single-trajectory agents—Claude Code and OpenHands—while costing less than either. Relative to its single-trajectory backbone, SWE-Xplorer improves resolved rate with all three evaluated models on both SWE-bench Verified and SWE-bench Multilingual. Together, these results show that iteration-based search can turn additional test-time compute into continued gains for long-horizon coding agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.