acceptodds
Under review as a conference paper at ICLR 2027

HarnessWeaver: Continual Self-Improvement through Task-Conditioned Harness Composition and Model Routing

Abstract

Continual self-improvement requires agents to use accumulated experience on subsequent tasks, yet failures may arise from limited executor capability or from inadequate planning, tool use, and execution control in the harness. We introduce HarnessWeaver, a two-timescale framework for continual self-improvement through task-conditioned model routing and harness composition without training model parameters. At task time, the Weaver uses accumulated evidence to select an executor and compatible components from a typed Living Pool, favoring low-cost configurations. Between rounds, it proposes model, harness, or joint revisions according to failure mechanisms and uses execution-based validation to admit persistent updates. On ALFWorld and Omni-MATH, frozen pre-adaptation references distinguish first-encounter transfer to new batches from recovery on repeated tasks. HarnessWeaver raises the equal-weight macro-average first-encounter success from 31.5% for the frozen reference to 56.8% (+25.3 points), outperforming the strongest external baseline, ACE at 39.1%, by 17.7 points. In independently rerun ALFWorld controls, Full reaches 58.6% first-encounter success, compared with 54.3% for Harness-only and 48.6% for Model-only, and also achieves the highest final-round success. These results show that reusable model–harness assemblies can serve as units of adaptation that connect accumulated experience, selective revision, and subsequent execution into a continual self-improvement process.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.