acceptodds
Under review as a conference paper at ICLR 2027

GP-Harness: Sequential Bayesian–Evolutionary Optimisation of LLM Harnesses

Abstract

Large language model agents rely on harnesses that organise planning, tool use, verification, and failure recovery around a fixed model. Automatic harness search can generate more candidates than can be evaluated on long-horizon tasks, making the choice of which candidates to evaluate a central challenge. We introduce , which combines evolutionary candidate generation with Bayesian optimisation to guide this choice. Starting from eight adapted open-source harnesses, programmatic mutation and recombination generate candidate pools. A Gaussian-process surrogate over graph features uses previous evaluations to predict candidate fitness and uncertainty, and expected improvement selects one candidate for evaluation in each round. Within each benchmark, all harnesses use a shared executor, with the task model, tool interfaces, evaluator, and per-task execution limits held fixed. With 20 new-candidate evaluations per run, achieves 64.6% mean success on a held-out 24-task Terminal-Bench 2.0 subset and a 43.3% mean resolved rate on the official 300-instance SWE-bench Lite test set across six search seeds, including runs that fall back to an initial harness. These results improve on the initial harness selected by its performance on the search tasks by 10.4 and 8.0 percentage points, respectively, and on budget-matched surrogate-free evolution by 4.9 and 4.6 points. Online ablations favour graph features over dimension-matched text features and uncertainty-aware acquisition over posterior-mean selection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.