acceptodds
Under review as a conference paper at ICLR 2027

Joint Model-Harness Optimization: Cost-Aware Resource Allocation for LLM Agents

Abstract

An LLM agent's quality and cost depend jointly on its model and harness, yet execution-time gains alone do not establish whether harness optimization is cost-effective. We study model–harness design as resource allocation, accounting separately for the cost of searching for a harness and the cost of executing it. We introduce HARP-Static, which combines component addition, trace-guided graph rewrites, and backward compression while tuning node-level resources and model assignments. Across six benchmarks, with harness search on five, we find both opportunities and limits to reallocating computation. Two searched architectures reduce to a minimal reasoning chain, and development-set analyses identify verification changes that lower cost without changing observed scores; larger savings from cheaper node-level models incur quality loss. On HotpotQA, Haiku-4.5 with a searched harness outperforms Opus-4.5 with its default harness by 10.4 percentage points at 55.3% of its execution cost. However, the only finite break-even point estimate in the matched-quality analysis requires approximately 1.72 million served tasks, and the underlying saving is not statistically established. Generalization poses a further constraint: on BFCL, the compressed harness loses 15.0 percentage points to the official harness, with post-hoc analysis locating the entire deficit in multi-turn tasks requiring repeated tool interaction. These results show that evaluating harness optimization requires distinguishing local execution gains from deployment-level savings and testing whether compression preserves the interactions a task requires.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.