acceptodds
Under review as a conference paper at ICLR 2027

EvoHarness-RL: Learning Runtime Harness Coordination for Self-Evolving Agents

Abstract

Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, recover from failures, and reuse experience across extended interactions. Yet existing harnesses and their use are often tailored to environments and controlled through prompts, heuristics, or system-specific rules, making agent and harness coordination difficult to jointly optimize. We introduce EVOHARNESS-RL, a unified framework that separates environment-specific harness implementations from a shared policy-facing interface. EVOHARNESS-RL organizes external support into a Belief, Progress, and Experience (BPE) workspace and exposes four compact harness actions for accessing and updating this state. We first instantiate BPE as an inference-time scaffold and then make harness coordination learnable through supervised initialization followed by cost-aware GRPO. Across heterogeneous long-horizon tasks, EVOHARNESS-BASE improves the average success rate of frontier models by 10.0 percentage points, while EVOHARNESS-RL outperforms the strongest open-source baseline by 8.5 percentage points, together with higher RL rollout efficiency and stronger generalization to unseen tasks. Our analyses show that training gradually shifts agents from frequent scaffold use toward selective, environment-dependent harness access as the policy becomes more capable, while the external workspace continues to evolve and refine itself to better support task execution and generalization. Together, these results show that long-horizon agents benefit not only from external scaffolding itself, but also from learning how and when to coordinate with external support as part of a cost-aware policy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.