HOP: Harness Optimization with Provable Guarantees
Abstract
A language-model agent's harness governs its reasoning and tool use, but its effects on performance depend on complex task interactions. This complexity often leads harness optimization to be treated as a black-box problem. We take a novel view of harness optimization by distinguishing structured harness configurations from unstructured task contexts. Building on this distinction, we introduce HOP, a semiparametric framework that models outcomes parametrically in configurations and nonparametrically in task contexts. This structure allows HOP to estimate performance across an exponential configuration space while testing only a linear number of configurations, substantially reducing the data collection burden. Under suitable conditions, HOP's outcome estimators are consistent, asymptotically normal, and asymptotically efficient, and HOP yields consistently optimal harness selection. On -Bench (Shi et al., 2026), HOP achieves lower regret and fewer constraint violations than leading harness optimization methods under a common evaluation protocol.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.