RLVR from Within: Learning to Ask to Elicit Intrinsic Rewards for Test-Time Reasoning
Abstract
Test-time reasoning for large language models (LLMs) often relies on external process reward models (PRMs) or value estimators to guide search over intermediate reasoning states. However, learning such rewards from final outcomes poses a challenging credit-assignment problem, and the reasoner’s own self-verification may not reliably predict final-answer correctness. To address both challenges, we propose AskR, which learns what to ask to make the reasoner’s self-verification more predictive of final correctness and thus more useful as an intrinsic reward. Rather than learning an external PRM, AskR learns a sub-question policy that structures reasoning into sequences of sub-questions and sub-answers (SubQAs), exposing dense self-verification signals at each SubQA. By shaping which intermediate reasoning states are exposed, AskR makes the induced intrinsic reward more predictive of final-answer correctness and therefore more useful for test-time optimization. Across our benchmarks, AskR consistently improves over PRM-guided reasoning under the same optimizer, by 7.12 points on average and up to 9.18 points depending on the optimizer. In its strongest setting, AskR further outperforms Reinforcement Learning with Verifiable Rewards (RLVR) on average, while using only 5% and 17% as much training data as RLVR on text and visual reasoning tasks, respectively. Our results suggest a new route to test-time reasoning: instead of learning how to evaluate arbitrary reasoning states, learn which states to expose so that the reasoner's existing verification signal becomes useful.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.