ClawGym II: Exploring Black-Box Reinforcement Learning on Agent Harness
Abstract
Agent harnesses have substantially improved performance on long-horizon tasks by coordinating model interactions with complex environments. However, reinforcement learning through such harnesses remains underexplored, as scaling such training to long-horizon agent tasks introduces fundamental challenges. In this work, we present a unified black-box RL framework for stable and scalable optimization of general agents through complex harnesses. We first build a sandbox-based execution infrastructure that isolates task environments and native harnesses within temporary sandboxes for large-scale concurrent rollouts. We then decouple policy optimization from opaque harness execution and place a serving proxy at the model boundary to capture faithful model calls, reconstruct structured training trajectories, and preserve training–inference consistency through black-box token-in-token-out. Crucially, rather than training separate harness-specific policies, we formulate each task–harness pair as a basic training instance and jointly optimize a shared policy from rollouts produced by heterogeneous harnesses. With Qwen3-30A3B jointly trained through OpenClaw and Claude Code, our model improves Pass@1 on ClawGym-Bench by points under OpenClaw and points under Claude Code, while improving PinchBench by and points, respectively. The same framework also delivers consistent gains on more challenging task settings from JobBench and OfficeQA. Overall, our results show that heterogeneous black-box harnesses can be jointly used as training interfaces for a single general agent policy, enabling effective and scalable optimization across diverse execution systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.