acceptodds
Under review as a conference paper at ICLR 2027

Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable

Abstract

The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, execution environments, and application requirements change, the harness must be continually modified to add capabilities or adapt existing behaviors. Before a human developer or coding agent can make such a change, they must identify all code locations that implement the target behavior. This is difficult because production harnesses are often large, tightly coupled, and behaviorally distributed across files, functions, execution stages, and state transitions. Modification requests describe what the system should do, whereas repositories are organized by files, functions, and modules. They do not directly reveal the complete implementation path of a behavior. Existing approaches to code search, repository indexing, and long-context processing make code easier to inspect, but they still leave developers and coding agents to recover this mapping themselves. Behavior localization is therefore a central bottleneck in harness evolution. We introduce the Harness Handbook, a behavior-centric representation synthesized automatically from a harness codebase through static program analysis and LLM-assisted behavioral structuring. The Handbook organizes implementation knowledge around system behaviors and links each behavior to the corresponding source code. We also introduce Behavior-Guided Progressive Disclosure (BGPD), which guides coding agents from high-level behavior descriptions to relevant implementation details and verifies candidate locations against the current source. We evaluate Harness Handbook on diverse modification requests from two open-source agent harnesses by comparing planning with and without Handbook access. Handbook-Assisted planning improves behavior localization and edit-plan quality while using fewer planner tokens. The largest gains appear for changes involving scattered implementation sites, rarely executed code paths, and cross-module interactions. These findings indicate that evolving complex agentic systems depends not only on generating edits, but also on determining where those edits should be made.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.