acceptodds
Under review as a conference paper at ICLR 2027

Self-HPO: Self-Evolving Coding Agents via Harness-Policy Optimization

Abstract

While large language model (LLM) agents have shown promising potential in software engineering, their performance is jointly determined by the model weights and the harness that mediates their interaction with the environment. The two are mutually reinforcing: a better harness exposes and compensates for the weaknesses of the current weights, while stronger weights allow the harness to elicit more capable behavior. Yet existing methods optimize one component with the other fixed, so a harness tuned for one checkpoint may no longer suit an updated policy, while policy training remains limited to the trajectories that its fixed harness elicits. In this work, we propose Self-HPO, a recursively self-improving framework in which the harness and the policy co-evolve through alternating optimization. Specifically, in each round, the current policy first analyzes its own failures and proposes harness revisions (H-step); reinforcement learning (RL) then updates the policy under the selected harness (P-step), and hard tasks encountered during training guide the next H-step, where the updated policy serves as the proposer. Experiments show that Self-HPO consistently outperforms harness-only and policy-only optimization, improving Qwen3.5-35B-A3B from 61.0% to 69.5% on SWE-Bench Verified, with gains extending to out-of-domain benchmarks (e.g., 54.2% → 59.7% on SWE-Bench Multilingual). Extensive analysis indicates that the gains are retained in both components and are complementary, and that fewer training prompts fail on every rollout during co-evolution than under RL with a fixed harness. Finally, we show that with a stronger external model as the harness proposer, the same framework yields a further performance boost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.