acceptodds
Under review as a conference paper at ICLR 2027

DuetHarness: Bidirectional Policy–Harness Co-Optimization for LLM Agents

Abstract

The capability of a large language model (LLM) agent depends jointly on its policy and its harness, the executable runtime that mediates interaction with the environment. Existing work either optimizes one side or updates both, but learning still remains essentially one-sided. Co-optimization is difficult because the policy and harness control separate strategies, yet each update changes the problem faced by the other. In this paper, we cast their interaction as a common-payoff game in which each side is improved against the other's current configuration. The idealized fixed point of this coupling is a mutual best response, a Nash equilibrium of this game. We therefore introduce DuetHarness, which implements this coupling by alternating learned updates analogous to approximate best-response steps. With the policy frozen, the editor learns to revise the harness from rollout feedback; with the harness fixed, the policy adapts to it by training against both the current and historical harnesses. Across three agent benchmarks with Qwen3.5-9B, DuetHarness consistently outperforms the corresponding vanilla agent and strong baselines. Ablations show that these gains rely on the alternating co-optimization rather than simply optimizing both sides. Our code is available at https://anonymous.4open.science/r/DuetHarness_ano-F0CE/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.