ConSeqHarness: Memory-Conditioned Sequential Optimization of Agent Harnesses
Abstract
As large language models (LLMs) increasingly rely on tools, routing, persistent state, and verification for long-horizon execution, agent performance depends not only on the model itself but also on the external harness that organizes its execution. Existing optimizers use prior code, scores, and traces to guide new proposals, but typically treat such experience as proposal context rather than reusable, validated optimization state. We introduce the verified harness transition as a reusable optimization primitive that binds a motivating failure to a typed and scoped edit, its measured gain, validation evidence, and non-target effects. This turns sequential harness optimization into a memory-augmented process in which previously validated changes can explicitly condition future decisions. We formalize this view as the Harness-Reflected Decision Process and instantiate it with Memory-***Con***ditioned ***Seq***uential Optimization of Agent ***Harness***es (ConSeqHarness), a constrained sequential optimizer that retrieves compatible verified transitions, proposes bounded structural edits, and atomically updates both the harness and transition memory only after structural, runtime, improvement, and regression checks. Across three target models and five public agent benchmarks, ConSeqHarness ranks first in all 15 settings and improves macro-average accuracy over the strongest baseline by 2.73–3.43 points under the same candidate-attempt budget. Controlled ablations further show that retaining accepted experience and representing it as structured transitions both contribute to the gains, while unverified memory and unconstrained editing substantially increase regressions and runtime failures. *The source code is available at https://anonymous.4open.science/r/ConSeqHarness-579E.*
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.