acceptodds
Under review as a conference paper at ICLR 2027

GymHarness: Recursive Self-Improvement of Agents with Recursively Verified Verifiers

Abstract

A self-improving agent learns from data it generates itself, so how far it improves is decided by the verifier that labels that data. When ground truth is scarce, the verifier must be built by the system, and a fixed verifier does not stay accurate: as the policy trains, it leaves the distribution on which the verifier was calibrated, and the verifier accepts more of its errors. We present GymHarness, a recursive self-improvement framework in which a policy and a smaller verifier improve each other through supervised fine-tuning, and in which every verifier update is itself verified. A candidate verifier replaces the current one only if it passes a hidden, append-only probe suite whose labels come from execution evidence that the agent cannot write, so the recursion bottoms out in exogenous evidence rather than in the model’s own judgment. Trained only on its own verified trajectories, Qwen3.6-27B improves by 11.3 points on WebShop, ALFWorld, and DBBench and by 5.2 points on two held-out environments, recovering 89% of the gain of oracle-filtered training with fixed anchors and a 10% adaptive rollout-audit rate. A 9B verifier suffices to supervise the 27B policy, and co-evolving the verifier without the recursive gate performs worse than not evolving it at all.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.