acceptodds
Under review as a conference paper at ICLR 2027

Judging RLVR Boundary Comparisons Against Their Own Measured Noise

Abstract

Whether reinforcement learning with verifiable rewards (RLVR) expands a base model's reasoning boundary is disputed in print. We show that both published answers can be read off one fixed (base, RLVR) pair at every sampling budget we test by moving only sampling temperature and scoring rule; the raw grid counts behind this are not floor-judged, and the base-ahead ones concentrate in chat-template cells at low temperature. Every movement is then judged against measured noise. Of the 30 registered comparisons between the two temperatures this debate uses, 11 are unjudgeable, moving less than the noise floor of the cells they compare; the rest sit on two Minerva cell pairs, and the 11 of them reseeded all clear again. In the sharpest instance the crossover point k⋆ moves from 10 to 48 on the same problems, scoring rule and seed. A post hoc observation at k=1 gives the movement a mechanism: temperature lowers the base arm's pass@1 by a larger factor than the RLVR arm's, and the sign of that asymmetry matches the direction k⋆ moves in every cell where it is defined, though a constant rule matches all but one. We turn the measurement into a decision procedure that returns AHEAD, CAUGHT or UNJUDGEABLE for a pair under a protocol, with its false-AHEAD rate set by a null control of two draws from one arm; neither the measured noise floor nor this null-calibrated gate has a counterpart we found in this literature. On seven released pairs the verdict changes with temperature for two of them at the null-calibrated gate: the answer belongs to the protocol as much as to the pair, and it names a finite-budget advantage rather than a support relation. The harness, the noise-floor and lineage instruments, and judge_boundary.py are in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.