Harness Training: A Post-Training View of Agent Harness Evolution
Abstract
Large language model agents depend on an agent harness that comprises persistent prompts, skills, tools, middleware, and workflows governing their behavior. Harness evolution updates this state after deployment without changing model weights. Existing work is commonly categorized by the component being edited or the search algorithm, rather than by whether an update is retained to reproduce target behavior or improve execution outcomes. We frame harness evolution as Harness Training, an analogue of parameter post-training at the level of learning objectives. This view distinguishes Harness-SFT, which learns from target trajectories, from Harness-RL, which learns from reward feedback. Experiments show that Harness-SFT makes agent executions more closely follow target trajectories, while Harness-RL changes retained behavior in the direction predicted by the reward objective. When combined, Harness-SFT increases the availability of task-solving edits, and Harness-RL uses reward feedback to identify and retain them, improving reward search. Based on this finding, we develop Signal-Targeted Harness Training (STHT), which applies Harness-SFT to tasks failed by the initial harness before initializing Harness-RL from the supervised state. Across diverse experimental settings, STHT reaches validation thresholds with fewer rollouts and outperforms competing methods on unseen tasks. These gains transfer across harnesses, reward optimizers, and teachers. Code is available anonymously at https://anonymous.4open.science/r/harness-train-45C6/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.