acceptodds
Under review as a conference paper at ICLR 2027

JH-GRPO: Joint-Harness Group Relative Policy Optimization for Code Agent

Abstract

Coding agents can solve a repository task through one harness yet fail through another. Existing mixed-harness reinforcement learning optimizes mean reward rather than taskwise co-success across interfaces. We introduce Joint-Harness Group Relative Policy Optimization (JH-GRPO), which reallocates a fixed-size comparison group across native harnesses for the same task, without additional rollouts. To exploit this joint supervision, JH-GRPO optimizes for cross-harness co-success while assigning trajectory-specific credit and preserving harness-specific tool-use constraints. Specifically, a Joint Harness Coverage Utility rewards higher-order co-success, a Marginal Joint Advantage assigns this nonlinear utility to individual trajectories, and harness-conditional constraints localize penalties for invalid and foreign-schema tool calls. We also introduce Joint Harness Reliability (JHR) to measure taskwise co-success beyond mean Pass@1. Evaluated with Qwen3.5-35B-A3B across four coding harnesses and four software engineering benchmarks, our JH-GRPO improves macro Pass@1 from 48.8% to 53.7% and JHR@4 from 24.6% to 35.8% over the strongest mixed-harness baseline under matched compute. It improves worst-harness and unseen-harness Pass@1 by 6.6 and 3.5 percentage points, respectively, while reducing foreign-schema violations by 46.4%. These results demonstrate that taskwise joint optimization effectively improves cross-harness reliability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.