acceptodds
Under review as a conference paper at ICLR 2027

Divide and Evolve:Orchestrating Agent Harness Evolution with Diagnostic Multi-Agent System

Abstract

LLM-driven agents increasingly rely on well-designed harnesses to execute complex long-horizon tasks, yet harness design flaws can recur across tasks and require costly human intervention to resolve. Harness evolution addresses this by automatically searching for and implementing harness modifications based on execution traces and task feedback, but effective evolution demands sound multi-stage decision coordination and principled control of modification boundaries and search directions, challenges that existing approaches fail to address simultaneously. To this end, we propose DiVE, a diagnosis-driven multi-agent harness evolution framework that integrates constraint-based boundary control and experience-guided search. DiVE assigns dedicated roles to harness analysis, repair planning, and patch implementation, propagating decision rationale through explicit intermediate outputs to support error localization, while incorporating program constraints, historical experience, and task feedback to jointly govern modification scope and search direction across iterations. Extensive experiments on -Bench and Terminal-Bench demonstrate that DiVE consistently improves harness performance across diverse task environments, and that small LLMs trained on complete evolution trajectories can also produce effective harness improvements, suggesting that structured evolution trajectories provide useful supervision for learning evolution capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.