EvoRule Harness: Lifelong Harness Adaptation with Validated Intervention Rules
Abstract
As general-purpose models are increasingly deployed with fixed weights, harness optimization is emerging as a key paradigm for adapting fixed-weight models to diverse tasks and deployment settings by managing context, tool use, and memory. However, deployment is inherently non-stationary, as tasks, tools, and interaction patterns change over time, while existing memory mechanisms are often rigid, training-dependent, and brittle, making reliable evolution over long horizons challenging. To address these limitations, we propose *EvoRule Harness*, a plug-and-play, continually adaptive memory system for long-horizon harness evolution. Given a frozen backbone model and an existing harness optimizer, EvoRule Harness builds a dynamic two-tier memory of causal rules. Each rule specifies an application condition and an intervention, and tracks execution-outcome feedback. Rules are revised online through individually traceable patches, allowing the effect of each update to be tracked and rolled back when necessary. Rule updates first enter a volatile tier and are promoted to stable memory only after repeated validation across rounds and data subsets, while unreliable updates are discarded. EvoRule Harness can be seamlessly integrated with diverse existing harness optimization methods, without requiring gradients or additional training. We evaluate EvoRule Harness on diverse tasks, including retrieval-augmented mathematical reasoning and interactive agentic tasks. EvoRule Harness improves Meta-Harness, TTHE, and Continual Harness by 7.0%, 8.5%, and 35.7% in average over Tau-Bench, Terminal-Bench-2, and SWE-Bench Verified, respectively. In particular, Meta-Harness augmented with EvoRule Harness achieves the best overall average of 67.4% among all harness optimization framework baselines, improving the base Meta-Harness from 63% by 7.0%. Our code is available at https://anonymous.4open.science/r/EvoMem-2C26/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.