acceptodds
Under review as a conference paper at ICLR 2027

How good is your harness? Statistical evaluation of coding harnesses

Abstract

AI agent benchmarks conflate the underlying language model with the harness that wraps the model in tools, prompts, and control flow. We develop a simple statistical method to disentangle the effects of the LLM and its harness on the agent's score; i.e., to attribute the variation in agents' scores to their LLMs and harnesses. We use the method to evaluate harnesses and LLMs on Terminal-Bench 2.0, and our results show that 1. the harness can matter as much as an LLM upgrade: the range of harness effects is comparable to the gain from Claude Opus 4.1 to Claude Opus 4.6. 2. the harness effects can be heterogeneous; i.e., some harnesses work better with some LLMs than with others. 3. the effects of individual harness features can also vary across LLMs. Our results confirm and, more importantly, quantify practitioners' intuition on the importance of the harness and validate efforts to tailor harnesses to specific applications.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.