acceptodds
Under review as a conference paper at ICLR 2027

Where to Think, What to Explore: Hierarchical Agentic Reasoning and Exploration for Coding Agents

Abstract

Coding agents work by interacting with a repository over many turns, so the compute they spend at inference time is as much a design choice as the model itself. That compute can go into reasoning inside a decision, more turns, more candidate actions, or parallel subagents, and these choices have not been compared under one accounting. We measure the four dimensions (reasoning depth , interaction horizon , action width , subagent fan-out ) and their interactions under one coding agent, model, benchmark, serving stack, and cost accounting. Scaled alone they saturate for different reasons: depth has a finite operating point, horizon reaches diminishing returns, width produces candidate diversity that naive selection fails to exploit, and delegation uptake limits passive subagents. They also interact: reasoning changes both the value of later interaction and the cost of parallel branching, and the depth of the step that reduces a fan-out to one action sets what that fan-out is worth. These measurements motivate HARE, Hierarchical Agentic Reasoning and Exploration, which treats and as a serial pair and and as a parallel pair. HARE reasons only at the turns that write to the repository or synthesize gathered evidence, and folds each fan-out back into one committed action with a single reduction step. On SWE-bench-Pro tasks with a Qwen3.8-27B model, HARE matches dense reasoning at 15% less wall-clock and raises the best measured resolve rate from 56.7 to 60.0%. Whole-trajectory sampling and sequential refinement stay below the frontier HARE reaches, at higher cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.