HarnessPro: Recursive Harness Self-Improvement with Pareto Optimization for Performance–Cost Trade-offs
Abstract
Harness development plays a crucial role in building capable AI agents. Recent approaches improve harnesses through recursive self-improvement (RSI), but their performance gains often come with higher execution costs. We introduce HarnessPro, a label-free method for evolving harnesses to improve task performance while reducing execution cost. In this setting, candidate harnesses are evaluated using LLM-generated tests without access to ground-truth solutions or official tests. Specialized agents diagnose execution trajectories, plan updates with different performance-cost priorities, and implement them as candidate harnesses. Rather than returning a single harness, HarnessPro uses normalized weighted squared-distance selection to provide cost-focused (lite), balanced (medium), and performance-focused (heavy) harnesses for different deployment budgets. Selected harnesses enter the next iteration, where execution feedback guides further updates. Our experiments cover Terminal-Bench 2, SWE-Bench Pro, and ProgramBench, with GPT-5.5 and Codex as well as GLM-5.2 and Claude Code. Across these benchmarks with GPT-5.5 and Codex, the heavy mode of HarnessPro averages a 2.0 percentage point gain in pass rate and a 26.1% reduction in per-task cost relative to the strongest prior label-free baseline. This advantage also holds for GLM-5.2 and Claude Code. Code is available at [this link](https://anonymous.4open.science/r/HarnessPro-ICLR-CDB2).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.