GP-Harness: Sequential Bayesian–Evolutionary Optimisation of LLM Harnesses
Abstract
Large language model agents rely on harnesses that organise planning, tool use, verification, and failure recovery around a fixed model. Automatic harness search can generate more candidates than can be evaluated on long-horizon tasks, making the choice of which candidates to evaluate a central challenge. We introduce , which combines evolutionary candidate generation with Bayesian optimisation to guide this choice. Starting from eight adapted open-source harnesses, programmatic mutation and recombination generate candidate pools. A Gaussian-process surrogate over graph features uses previous evaluations to predict candidate fitness and uncertainty, and expected improvement selects one candidate for evaluation in each round. Within each benchmark, all harnesses use a shared executor, with the task model, tool interfaces, evaluator, and per-task execution limits held fixed. With 20 new-candidate evaluations per run, achieves 64.6% mean success on a held-out 24-task Terminal-Bench 2.0 subset and a 43.3% mean resolved rate on the official 300-instance SWE-bench Lite test set across six search seeds, including runs that fall back to an initial harness. These results improve on the initial harness selected by its performance on the search tasks by 10.4 and 8.0 percentage points, respectively, and on budget-matched surrogate-free evolution by 4.9 and 4.6 points. Online ablations favour graph features over dimension-matched text features and uncertainty-aware acquisition over posterior-mean selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.