Hypothesis-Conditioned Evaluation for Efficient Harness Optimization
Abstract
Harness optimization improves an LLM agent by editing its prompts, tools and helper code with model weights fixed. Evaluating every edit on every training task is expensive, but a small random sample can miss the failure that motivated the edit. We introduce Hypothesis-Conditioned Evaluation (\method), which uses the editor's explanation to guide evaluation. Before any candidate rollout, the editor declares an evaluation contract: the failure it targets, the tasks it expects to fix or might break, an activation predicate specifying where the edit applies, and a runtime event showing that it fired. Conditional prompt rules are inserted only when their predicates hold. The contract's targets and risks are always evaluated; a stratified probability sample fills the remaining budget, including tasks outside the declared scope. An augmented inverse-probability-weighted estimator combines accepted task history with sampled outcomes to compare candidates on the full training-score scale. Retention requires a higher estimated score, observed firing within scope, and improvement on a named target where the edit fired. In two ten-edit comparisons with Agentic Harness Engineering (AHE), using GPT-5.4, used \EvaluationTokenSaving% fewer scored search-evaluation tokens on Terminal-Bench (30 training, 58 held-out tasks) and \SoftwareEvaluationTokenSaving% fewer on Software100 (33 training, 67 held-out). The training-selected harnesses achieved held-out three-attempt mean success rates () of \HCEFinal% versus \AHEFinal% on Terminal-Bench, and \SoftwareHCETestScore% versus \SoftwareAHETestScore% on Software100. Earlier SWE-bench studies also used fewer evaluations with matching held-out quality. The reported searches establish evaluation savings, but do not establish general quality superiority or isolate the benefit of hypothesis-based task selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.