Harness-Bench: Evaluating Agent Harnesses across Models and Realistic Workflows
Abstract
Agent harnesses shape how language models turn instructions into completed work by governing tool use, context management, and execution. Choosing a harness therefore requires understanding what different harnesses enable the same model to deliver, and at what cost. We introduce **Harness-Bench**, a benchmark and reusable evaluation pipeline for comparing native harness configurations on realistic work tasks. The suite comprises 300 sandboxed tasks constructed through log-based and domain-based synthesis, with explicit input resources, deliverables, and operational constraints. Task-specific oracles score final artifacts, awarding partial credit for completed work modules under safety and risk gates. The pipeline fixes external task conditions and scoring rules, preserves native harness behavior, and records token usage and API expenditure. We evaluate seven harnesses with ten model backends on general tasks and seven on multimodal tasks, using paired comparisons on shared tasks. Our evaluation reveals substantial, model-dependent performance differences: no harness performs best across all ten backends, and GLM-5.3's general-task full-credit rate varies by 13.5 percentage points across harnesses. Harness choices also change with the delivery objective: configurations that lead mean score need not lead full-credit attainment. Higher expenditure does not ensure better delivery: with GPT-5.6 Sol, Nanobot incurs 2.97 times OpenClaw's API cost while achieving a lower full-credit rate. Stratified trajectory analysis and paired cases explain concrete delivery gaps: corrected tool calls can leave workflows unfinished, and mutually consistent outputs can still omit required evidence. These findings connect model-dependent harness performance to the execution steps that turn partial progress into delivered work. Harness-Bench brings task outcomes, unmet requirements, and resource costs into a common evaluation framework, providing an empirical foundation for selecting and improving agent harnesses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.