ReHarness: Learning Reusable Harness Evolution Across Tasks
Abstract
Large Language Model (LLM) agents increasingly rely on programmable harnesses to coordinate reasoning, tool use, and execution. Existing harness evolution methods typically optimize a task-specific harness from scratch for each new task, requiring substantial task-specific data and evolution steps. In this work, we identify that harness evolution exhibits reusable structure across heterogeneous tasks: independently evolved harnesses repeatedly discover similar improvement strategies, although the resulting harnesses remain task-specific. Building on this insight, we propose ReHarness, a reusable harness evolution framework that learns how to evolve across tasks. ReHarness learns a transferable evolution strategy from multiple offline-given tasks and applies it to rapidly adapt task-specific worker harnesses to new tasks using only a few examples. Experiments across four LLMs and six benchmarks demonstrate that ReHarness consistently improves agent performance on four new tasks not used during the offline phase, while outperforming the task-specific evolution method Self-Harness on two tasks from the offline suite under the same data budget. Code and models will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.