acceptodds
Under review as a conference paper at ICLR 2027

Beyond Failed Trajectories: Diagnosing and Repairing LLM Agent Harnesses Across Three Observability Levels

Abstract

Failure-driven debugging of LLM-agent harnesses implicitly assumes that the execution selected for diagnosis already contains evidence of the defect. We show that successful executions violate this assumption in two distinct ways. In L2, the task succeeds but the trajectory contains abnormal evidence; in L3, the ordinary execution remains healthy until another reachable condition exposes the defect. Across 9,058 harness-related issues and merged pull requests from 16 open-source agent projects, 25.1% of records with defined levels are L2 or L3, with high agreement in independent audits by two authors. A controlled study of 36 reconstructed historical defects further shows that passive source inspection can raise suspicion about latent defects without confirming their manifestation, and that source localization remains a separate bottleneck once an exposing execution is available. We introduce Proactive Invariant Harness Tomography (Proactive IHT), which audits all trajectories, routes successful runs with invariant violations to diagnosis, selectively probes healthy candidates for missing evidence, and verifies candidate mechanisms before scoped repair. Across four public benchmarks and four model backbones, Proactive IHT improves over HarnessFix in all 16 matched settings by 2.0–5.0 percentage points in task completion rate (TCR), with an average gain of 3.6 points. These results show that harness repair benefits from treating evidence acquisition as part of the debugging problem rather than using task failure as its only entry point.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.