acceptodds
Under review as a conference paper at ICLR 2027

HOP: Harness Optimization with Provable Guarantees

Abstract

A language-model agent's harness governs its reasoning and tool use, but its effects on performance depend on complex task interactions. This complexity often leads harness optimization to be treated as a black-box problem. We take a novel view of harness optimization by distinguishing structured harness configurations from unstructured task contexts. Building on this distinction, we introduce HOP, a semiparametric framework that models outcomes parametrically in configurations and nonparametrically in task contexts. This structure allows HOP to estimate performance across an exponential configuration space while testing only a linear number of configurations, substantially reducing the data collection burden. Under suitable conditions, HOP's outcome estimators are consistent, asymptotically normal, and asymptotically efficient, and HOP yields consistently optimal harness selection. On -Bench (Shi et al., 2026), HOP achieves lower regret and fewer constraint violations than leading harness optimization methods under a common evaluation protocol.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.