Decodable Need Not Be Reward-Reachable: A Controlled Study of Direct-Answer Learning in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) depends on both the answers a policy samples and the update that uses their rewards. Guaranteeing correct answers in training removes a coverage problem, but leaves open whether the update can learn an answer that supervision can teach. We isolate this distinction with execution-labeled Yes/No questions about runtime errors in short Python programs. On a list-index task, Qwen3-8B representations support linear label decoding, and same-data supervised fine-tuning (SFT) reaches margin AUC 0.927. Forcing both final answers into every group restores reward variation and a nonzero answer-logit signal, yet the tested centered update reaches only 0.558; the same intervention succeeds on dictionary lookup. Changing the loss or weighting rule improves index readouts under matched sampling. The primary exact-gradient comparison remains inconclusive, while a separate pre-specified error-focused rule learns the direct index answer in all three seeds and meets its local-sufficiency criterion. Reasoning provides a distinct route: strong Qwen answers are already expressed before task-specific training, while direct answers remain weak after reasoning-channel training. The results separate recoverable label information, answers supervision can teach, and answers a specified reward update learns.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.