acceptodds
Under review as a conference paper at ICLR 2027

One Policy, Many Harnesses: Reinforcement Learning for Transferable Harness Adaptation

Abstract

Reinforcement learning has improved the performance of coding agents on repository-level software engineering tasks. However, it remains unclear whether the improvement after RL reflects success on more distinct tasks or on the same tasks under more harnesses. We conduct eight training experiments using PPO, alongside one GRPO comparison. Together with the base model, these yield ten Qwen3.5-27B checkpoints evaluated in 40 checkpoint–harness combinations on 731 SWE-Bench Pro tasks, combination produce 87,720 evaluation sessions for Pass@3 evaluation. The key findings are: ① Task-level analysis shows that harness-specific gains primarily transfer from tasks the base model had already solved through another harness. The main gains increase success shared across harnesses without a comparable expansion in total observed task coverage. ② Gains can reach harnesses excluded from RL training, but transfer is direction-dependent: improvements in one harness can coexist with regressions in others. ③ Execution analysis links improved task success to more frequent delivery of nonempty patches, highlighting the importance of carrying edits through to an executable artifact. Together, these findings identify harness adaptation as a substantial component of the observed RL gains. They motivate training and evaluating coding agents for reliable execution across deployment harnesses, rather than treating a higher score in one harness as evidence of broader task-solving competence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.