acceptodds
Under review as a conference paper at ICLR 2027

Count the constraints, not the tests: verifier diversity governs reinforcement learning from execution feedback

Abstract

Reinforcement learning from verifiable rewards (RLVR) trains a code policy against a test suite, so what it can learn is bounded by the distinctions that suite draws. Practice summarises a suite by the number of tests it holds. If no sampled program separates two tests, they state the same constraint, and a reward computed from twenty of them still takes only two values. We measure the distinctions instead, grouping a suite's tests by the pass and fail column they produce over twenty programs from a fixed checkpoint, and counting the distinct columns, the effective constraint count. The gap is wide in practice: among generations that pass at least one test of a five test suite, three quarters pass all five, a share that barely moves across three different models. We then hold the tasks and the test count fixed and vary only whether the three tests the reward reads come from three clusters or from one, over tasks with held out and three seeds per arm. The pass@1 rises from to , with every diverse seed above every redundant seed, while pass@20 moves by points, so the effect is on reliability, not reach. The count needs only samples from a checkpoint that already exists, so a suite can be judged before compute is spent on it. Count the constraints, not the tests.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.