Choosing How Coding Agents Continue: A Study on BigCodeBench
Abstract
After producing a program, a coding agent can stop, revise it, request a review, or start again. We ask whether information from the first attempt helps choose among these actions. On an eligible subset of BigCodeBench, we train a routing policy on exploratory executions and freeze it before evaluation on unseen tasks. Compared with a fixed strategy, regeneration by the stronger model, the observed gain was 2.5 percentage points (95% CI: -1.0 to 5.5). It was not statistically conclusive and fell below the preregistered practical target. A post-hoc analysis locates the net gain in stopping and stronger-model review. That review action performed poorly across the full evaluation set, but produced no execution errors where the policy selected it. All its execution errors coincided with output-limit hits; review and regeneration used different effective output allowances. These descriptive associations do not establish why the selections worked or whether they will recur. We report state-acquisition, continuation and learning costs separately. The task-level comparison exposes differences hidden by aggregate action rankings while leaving the generalisable benefit of adaptive continuation unresolved.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.