DuetHarness: Bidirectional Policy–Harness Co-Optimization for LLM Agents
Abstract
The capability of a large language model (LLM) agent depends jointly on its policy and its harness, the executable runtime that mediates interaction with the environment. Existing work either optimizes one side or updates both, but learning still remains essentially one-sided. Co-optimization is difficult because the policy and harness control separate strategies, yet each update changes the problem faced by the other. In this paper, we cast their interaction as a common-payoff game in which each side is improved against the other's current configuration. The idealized fixed point of this coupling is a mutual best response, a Nash equilibrium of this game. We therefore introduce DuetHarness, which implements this coupling by alternating learned updates analogous to approximate best-response steps. With the policy frozen, the editor learns to revise the harness from rollout feedback; with the harness fixed, the policy adapts to it by training against both the current and historical harnesses. Across three agent benchmarks with Qwen3.5-9B, DuetHarness consistently outperforms the corresponding vanilla agent and strong baselines. Ablations show that these gains rely on the alternating co-optimization rather than simply optimizing both sides. Our code is available at https://anonymous.4open.science/r/DuetHarness_ano-F0CE/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.