How Much, Not Where: RLVR's Capability Loss Is Dose-Controlled and Deep-Tail Only
Abstract
Reinforcement learning with verifiable rewards (RLVR) reliably improves pass@1, yet can shrink the set of problems a model solves at large sampling budgets. We ask where that damage lives, when it becomes visible, and what controls it. Instrumenting GRPO training of Qwen3-1.7B-Base on MATH and MBPP, we track a forward-only signal: the per-prompt decay rate λ̂ of the likelihood of frozen base-correct trajectories on boundary prompts, those the base model solves only rarely. We report three findings. The damage is invisible at standard sampling depths. Across our entire dose grid, pass@k stays significantly above base for every k ≤ 128 and entropy settles on a floor rather than collapsing; harm appears only beyond k ≈ 192, reaching −0.034 (95% CI [−0.057, −0.010]) at pass@256. Stopping at pass@64, the deepest endpoint in most of the inversion literature, would have shown no harm at all. The damage is dose-controlled, and the dose is measurable early. It grows monotonically with λ̂ × steps; λ̂ measured in the first 17% of training predicts the rate over the disjoint remainder (r = 0.79–0.94 on math) and grows superlinearly in learning rate. Rehearsal removes the damage in proportion to dose, not targeting. A dose- and schedule-matched control that anchors randomly chosen prompts protects as well as our signal-triggered scheme, and protection appears on prompts that were never anchored. The deliverable is a dose–response frontier trading deep-tail coverage against pass@1, not a per-prompt targeting rule. Code is available at https://anonymous.4open.science/r/rlvr-dose.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.