acceptodds
Under review as a conference paper at ICLR 2027

Lego-RL: Harness-Native Reinforcement Learning for Coding Agents

Abstract

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train–inference discrepancies decouple rollout behavior from policy updates. To address this, we present **Lego-RL**, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. Lego-RL is built upon three pillars:**(1) faithful optimization** via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; **(2) reliable execution** via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and **(3) observable training** through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. With GSPO, **LEGO-RL** improves Qwen3.5-35B-A3B's SWE-bench Verified resolve rates by , , and points under OpenHands SDK, Claude Code, and OpenCode, respectively, with rollout–training sampled-token probability correlations of at least . The gains generalize across harnesses; mixed-harness training and alternative RL algorithms further demonstrate the flexibility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.