HarnessOpt: Agent Optimization through Harness Kernels
Abstract
Optimizing a language-model agent requires coordinating its workflow, tools, and supporting components. Direct code editing is expressive, but patch size poorly captures the behavioral scope of an update. We introduce HARNESSOPT, a framework that optimizes a persistent Harness Kernel: a compact, structured specification of global workflow and local components. HarnessOpt scales SkillOpt-style refinement to complete harnesses by aggregating execution evidence, adopting bounded behavioral edits, and compiling each revised Kernel into coordinated runtime changes. Memory and patient selection guide iterative improvement while model parameters remain fixed. The same process supports black-box adaptation, white-box optimization, and from-scratch generation for individual task distributions or task mixtures. Evaluations with Codex, Claude Code, Pi, and standalone runtimes demonstrate strong task performance and successful transfer of learned harnesses across models, hosts, and benchmarks. Representation and execution studies show that Kernel editing can yield higher task success with substantially less host-code growth than direct runtime editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.