From Failures to Evolution: Behavioral Rubrics for Agent Harness Evolution
Abstract
Large language model (LLM) agents increasingly rely on harnesses that manage context and execution strategies beyond the underlying model. Harness evolution provides a promising approach to improving agent capabilities without updating model parameters. Existing approaches increasingly leverage execution evidence and structured feedback to improve agents. However, they usually lack an intermediate behavioral specification that separates **what the agent should do** from **how the harness realizes it**. Such a specification should capture desired behaviors, applicable contexts, and boundaries while remaining independent of specific harness implementations. We propose **rubric-guided harness evolution (RHE)**, which introduces **behavioral rubrics** as an implementation-agnostic interface between execution evidence and harness modifications. We define behavioral rubrics as specifications of desired agent behaviors, applicable contexts, and boundary constraints. By separating behavioral objectives from harness realizations, rubrics enable different implementations of the same behavioral requirement. Hierarchical experience memory further preserves the relationships among execution evidence, rubrics, and historical realizations, enabling reuse across evolution iterations. We evaluate our approach on **SWE-bench Verified** and **AppWorld**. Across three backbone models and two benchmarks, RHE consistently improves agent performance over existing agent improvement approaches. Cross-environment transfer experiments further show that behavioral rubrics discovered from SWE-bench Verified remain effective when transferred to AppWorld. These results demonstrate that behavioral rubrics enable generalizable harness evolution beyond trajectory-specific optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.