GymHarness: Recursive Self-Improvement of Agents with Recursively Verified Verifiers
Abstract
A self-improving agent learns from data it generates itself, so how far it improves is decided by the verifier that labels that data. When ground truth is scarce, the verifier must be built by the system, and a fixed verifier does not stay accurate: as the policy trains, it leaves the distribution on which the verifier was calibrated, and the verifier accepts more of its errors. We present GymHarness, a recursive self-improvement framework in which a policy and a smaller verifier improve each other through supervised fine-tuning, and in which every verifier update is itself verified. A candidate verifier replaces the current one only if it passes a hidden, append-only probe suite whose labels come from execution evidence that the agent cannot write, so the recursion bottoms out in exogenous evidence rather than in the model’s own judgment. Trained only on its own verified trajectories, Qwen3.6-27B improves by 11.3 points on WebShop, ALFWorld, and DBBench and by 5.2 points on two held-out environments, recovering 89% of the gain of oracle-filtered training with fixed anchors and a 10% adaptive rollout-audit rate. A 9B verifier suffices to supervise the 27B policy, and co-evolving the verifier without the recursive gate performs worse than not evolving it at all.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.