acceptodds
Under review as a conference paper at ICLR 2027

Dose–Response of Verifier Flaws in Synthetic Agentic RL Environments

Abstract

Agentic reinforcement learning increasingly trains on thousands of synthesized environments whose reward code is itself machine-written. We ask what happens to the trained policy when some fraction of those verifiers are wrong. Auditing the 1,000 open Agent World Model environments, we find that 16.8% of their 10,000 code verifiers pass when the agent does nothing, because the generated sample data already satisfies the task. We then run a controlled dose–response study. We replace the verifiers of 0 to 50% of training environments with over-permissive versions that also accept an exploit, keep the originals as ground truth, and train Qwen3-8B with group-relative policy optimization. Reported reward rises steadily with dose. True capability holds until half the pool is flawed, then drops by 7 points. What decides the damage is whether the policy must discover the exploit. Flaws that the untrained policy already satisfies inflate reward without changing what it learns. A latent flaw, a status flag the agent controls, is found only when enough environments pay for it; once found, it is used everywhere, including on held-out clean environments, where the policy sets the flag on over 90% of the tasks it fails. Across all arms and two model sizes, how far the exploit saturated during training predicts this transfer, which does not extend to an unrelated domain. Two cheap detectors catch most flaws: a zero-action probe run before training, which flags every flaw that pays without action and no clean environment, and per-environment reward telemetry during training (AUROC 0.8–0.9).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.