ASSAY: RECOVERING FILTERED ENVIRONMENTS WITH VERIFIER-GROUNDED ACTOR–CRITIC RL
Abstract
Code-derived reinforcement-learning environments are often selected using a screening model's rollout outcomes. Such selection mixes two questions: whether an environment is reliable, and whether a particular training recipe can use it effectively. We study this distinction in , where an all-pass/all-fail screen removes of tasks that have passed execution-consistency, leakage, and verifier-agreement checks. A unanimous training group gives GRPO zero reward-driven advantage, but a unanimous screening group does not establish that future training groups will be degenerate. We introduce , which re-admits these tasks, uses single-rollout asynchronous actor–critic optimization, and pretrains the critic on outcome-labeled records retained during environment construction. On the retained RL task pool, improves MiMo-V2.5 from GRPO's to on SWE-bench Pro and from to on DeepSWE. Expanding the pool raises these scores to and . The same expansion changes Val by percentage points under and under the evaluated GRPO configuration. Results on GLM-5.3-Flash show the same direction of improvement. Ablations, prompt-exposure controls, and training-cost measurements characterize the roles of value initialization, sampling, and asynchronous execution. These results support a practical conclusion: rollout screening should not be treated as an optimizer-independent verdict on training-data value.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.