The Real Bottleneck Is Not Reward Noise: Gradient Signal Degeneracy in Long-Horizon Agent RL
Abstract
Reinforcement learning with verifiable rewards depends on an automatic verifier, and on single-turn tasks a considerable amount of verifier error has proven tolerable: noise rates of up to 15 percent leave peak accuracy within two points of a clean baseline. However, a long-horizon trajectory receives a single terminal reward while giving a faulty verifier dozens of opportunities to be wrong about it, so that tolerating the same rate per trajectory would demand a per-turn rate an order of magnitude lower, and long-horizon training should on this reasoning be the more fragile of the two. Contrary to that expectation, we measure no accuracy cost in a multi-turn search agent at false-positive rates up to 0.5, across a fourfold range of realized horizons and a sixfold range of training length. In stark contrast to what accuracy reports, however, the fraction of sampled groups contributing no reward-driven gradient rises from 0.57 to 1.00 over that same range, so that a policy whose batch contains no such gradient anywhere remains indistinguishable, on accuracy, from one trained against a clean verifier. Subsequently, we trace the discrepancy to verifier error that attaches to the task rather than to the answer and therefore removes whole groups at once, and we obtain a parameter-free account of the collapse which transfers across two environments, three model families, false negatives and KL regularization. Our work shows that a verifier cannot be certified by the performance of the policy it trains, and that the fraction of groups producing no reward-driven gradient, obtainable at no cost during training, should be monitored in its place.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.