acceptodds
Under review as a conference paper at ICLR 2027

AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

Abstract

Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct. This work presents AGENTBUG-SMITH, an automated harness bug reproduction approach that continuously discovers and reproduces real-world harness bugs from open-source agentic systems. Across different backbone LLMs, AGENTBUGSMITH consistently outperforms existing bug reproduction techniques designed for general software, achieving 10.67% - 27.56% higher success rates of reproducing harness bugs. By applying AGENTBUG-SMITH to open-source agentic systems in the wild, we construct LIVE-HARNESS-BENCH, a live and extensible benchmark that currently contains 200 reproducible harness bugs. We further demonstrate the utility of LIVE-HARNESS-BENCH through two downstream applications. First, we use LIVE-HARNESS-BENCH as the evaluation benchmark to systematically evaluate state-of-the-art software agents, revealing their limited capabilities in repairing real-world harness bugs. Second, we use LIVE-HARNESSBENCH as a knowledge base of real-world harness bug fixes, from which reusable repair skills can be distilled to improve existing software agents, increasing their harness-bug repair rates by 6.32%. Together, AGENTBUG-SMITH and LIVEHARNESS-BENCH establish a scalable foundation for continuously evaluating and improving software agents on harness bug repair, turning real-world agent failures into executable evaluation instances and reusable knowledge for harness improvement, thus contributing to the ultimate goal of recursively self-improving agents

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.