Monitoring Is Not Self-Improvement: An Evidence Ladder for Agent Harnesses
Abstract
Agent harnesses increasingly monitor execution traces, intervene during a run, and revise prompts, tools, or policies from experience. Evidence that any one of these mechanisms operates, however, does not establish that the agent has improved. We introduce an evidence ladder that separates three claims: measurement validity, runtime intervention efficacy, and persistent improvement on independent future tasks. Each level is expressed as an evidence contract with a distinct estimand, required observations, and invalid shortcuts. We instantiate the framework in HarnessEvolve through a task-paired study of a two-check monitoring package on all 89 Terminal-Bench 2.0 tasks, comprising 267 matched task-condition evaluations per arm. The monitored harness obtains 60/267 verified successes versus 61/267 for the static harness (difference -0.4 percentage points; 95% CI [-6.0,+5.2]), while the four-rule operational endpoint also remains near zero (29.5% versus 29.8%). Applying the ladder shows why these outcome estimates cannot be replaced by mechanism telemetry: the checks have unequal exposure, the dominant operational event is outside their target set, and deterministic trace rules remain proxies whose semantic interpretation is limited by human annotation. Exploratory update studies separately verify proposal, gating, and deployment without treating execution of that loop as evidence of future-task improvement. The result is an auditable protocol for matching self-improvement claims to the experiments needed to support them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.