acceptodds
Under review as a conference paper at ICLR 2027

STRATA: Preserving Rare Solution Strategies in Reinforcement Learning with Verifiable Rewards

Abstract

Reinforcement learning with verifiable rewards (RLVR) is usually audited with pass@k, which overlooks the diversity of correct solution strategies. On a controlled benchmark with exhaustively enumerated, generator-verified strategy sets, we measure rare-strategy survival and find a dose-response decoupling: as GRPO training grows from 0 to 14.9M supervised tokens, AIME24 pass@64 rises (.400 → .467) while survival falls overall (1.000 → .122). DPH-RL preserves pass@k but retains only .146 of rare strategies; WiSE-FT weight interpolation reaches .390. We introduce STRATA (STrategy-Targeted Rehearsal Anchoring): verified teacher traces are re-scored under the current policy every R steps, and cross-entropy rehearsal targets low-likelihood traces at risk of being forgotten. At 3M supervised tokens—one fifth of overtrained GRPO’s budget—STRATA matches its best AIME24 pass@64 (.467), ties best AIME25 (.433) and AMC23 (.925) in our twelve-condition matrix, and retains 2.6× standard GRPO’s rare strategies (.626 vs .244) and 5.1× overtrained GRPO’s (.626 vs .122). Five-seed results, matched-compute controls, selector validation, and a manually annotated natural-solution audit support the benefit across model families and evaluation regimes. We release all code, the benchmark, and the matched audit protocol for reuse.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.