STRATA: Preserving Rare Solution Strategies in Reinforcement Learning with Verifiable Rewards
Abstract
Reinforcement learning with verifiable rewards (RLVR) is usually audited with pass@k, which overlooks the diversity of correct solution strategies. On a controlled benchmark with exhaustively enumerated, generator-verified strategy sets, we measure rare-strategy survival and find a dose-response decoupling: as GRPO training grows from 0 to 14.9M supervised tokens, AIME24 pass@64 rises (.400 → .467) while survival falls overall (1.000 → .122). DPH-RL preserves pass@k but retains only .146 of rare strategies; WiSE-FT weight interpolation reaches .390. We introduce STRATA (STrategy-Targeted Rehearsal Anchoring): verified teacher traces are re-scored under the current policy every R steps, and cross-entropy rehearsal targets low-likelihood traces at risk of being forgotten. At 3M supervised tokens—one fifth of overtrained GRPO’s budget—STRATA matches its best AIME24 pass@64 (.467), ties best AIME25 (.433) and AMC23 (.925) in our twelve-condition matrix, and retains 2.6× standard GRPO’s rare strategies (.626 vs .244) and 5.1× overtrained GRPO’s (.626 vs .122). Five-seed results, matched-compute controls, selector validation, and a manually annotated natural-solution audit support the benefit across model families and evaluation regimes. We release all code, the benchmark, and the matched audit protocol for reuse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.