How good is your harness? Statistical evaluation of coding harnesses
Abstract
AI agent benchmarks conflate the underlying language model with the harness that wraps the model in tools, prompts, and control flow. We develop a simple statistical method to disentangle the effects of the LLM and its harness on the agent's score; i.e., to attribute the variation in agents' scores to their LLMs and harnesses. We use the method to evaluate harnesses and LLMs on Terminal-Bench 2.0, and our results show that 1. the harness can matter as much as an LLM upgrade: the range of harness effects is comparable to the gain from Claude Opus 4.1 to Claude Opus 4.6. 2. the harness effects can be heterogeneous; i.e., some harnesses work better with some LLMs than with others. 3. the effects of individual harness features can also vary across LLMs. Our results confirm and, more importantly, quantify practitioners' intuition on the importance of the harness and validate efforts to tailor harnesses to specific applications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.