When the Answer Extractor Is Part of RLVR Evaluation
Abstract
In closed-set multiple-choice RLVR, one answer extractor often assigns reward and reports test accuracy. A score change under this shared reader can reflect a better answer choice, a more readable answer format, or both. We audit this coupling with frozen outputs and an independent output-contract intervention. On 1,270 common MedQA items, the shared parser reports a +5.62 point training gain, while two independent model readers report −0.52 points (DeepSeek; 95% CI [−1.71, +0.66]) and −1.94 points (Llama; [−3.33, −0.63]). With the same model weights and a 512-token budget, retaining the explanation but requiring a terminal Answer: X changes the shared-parser gain from +5.87 to +0.47 points; the training-by-prompt interaction is −5.40 points (95% CI [−8.37, −2.10]). The answer marker appears in 97.8–99.6% of completions, although exact final-line compliance varies by checkpoint. These results do not determine semantic truth from any single reader. They show that the reported RLVR gain depends materially on the interface that reads the answer, and that the shared reader should not be the sole validator.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.