acceptodds
Under review as a conference paper at ICLR 2027

EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?

Abstract

Modern LLM-based agents rely on an evolving harness of tools, reusable skills, and specialist agents, yet existing evaluations largely assume this harness is fixed. We introduce EvoHarnessBench, a benchmark for systematically evaluating agents under controlled harness evolution along three axes: tools, skills, and agents. Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity in the task stream while keeping the harness fixed, EvoHarnessBench places non-stationarity in the externally supplied harness itself. The benchmark contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks, comprising 802 tasks, 520 tools, 42 skills, and 62 agents. We evaluate two complementary settings that capture the central challenges of harness evolution: deployment evaluation, which measures whether agents retain previously accessible competence as the harness expands, and self-evolving adap- tation evaluation, which tests whether accumulated experience remains useful as new capabilities are introduced. Our experiments reveal three persistent limitations in state-of-the-art agents. First, harness expansion alone can degrade performance on previously solved tasks, leading to harness-induced forgetting. Second, gains from self-evolving adaptation are inconsistent across stages, capability axes, and environments. Third, retention and adaptation can be in tension: preserving earlier competence does not necessarily improve adaptation to newly introduced capa- bilities, and vice versa. Together, these results establish harness evolution as a distinct and important challenge for building agents that can continuously exploit new capabilities while preserving previously effective behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.