Harness: Recursive Agent Harnessing for an Open World
Abstract
LLM agents' capabilities depend not only on their underlying models but also on their execution harnesses. Recursive self-improvement (RSI) offers a promising approach to automatically evolve these harnesses. However, deploying RSI in open-world settings still faces two critical challenges. First, open-world tasks are typically unseen during the harness design phase and novel at test time, leaving no ground-truth verifiers to reliably guide online harness edits. Second, the RSI process remains highly task-specific and fails to generalize to novel scenarios. We introduce Harness, a plug-and-play *recursive agent harnessing* framework that addresses these challenges through two connected levels of recursion. *Task-level recursion* refines the current task's harness using a contrastive Propose–Probe–Compose step that reliably guides edits without ground-truth verifiers and can be scaled further via parallel or sequential recursion. *Harnessing-level recursion* then extracts experience from these edits to update a continual improvement harness, enabling the agent to better adapt to future tasks. Across three professional benchmarks and four model–harness configurations, a single step consistently improves GLM-5 and Gemini 3.5 Flash over their base harnesses, yielding gains of up to 11.5 points on WorkBuddy-Bench, 4.5 points on JobBench, and 11.8 percentage points on LAB, with further gains as the task-level recursion budget increases.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.