Can RLVR Models Think More Creatively? A Study of Overconfidence and High Entropy Segment
Abstract
Recent advances in Large Language Models (LLMs) are often attributed to improved reasoning abilities through Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR). Although RLVR models perform strongly on mathematical benchmarks, their ability to generate creative solutions remains an open question. In this work, we compare RLVR models and math-pretrained models from the DeepSeek and Qwen families on mathematical problem solving with supplied reference solutions. RLVR models yield fewer responses judged both correct and novel, while their self-scoring high lower-tail token likelihood (HL) rates exceed those of the Math models by 15.85 and 74.70 percentage points, respectively. HL operationalizes token-level Overconfidence (OC) in this study, rather than calibration error. Motivated by the hypothesis that SFT and policy optimization reshape token probabilities, we analyze High Entropy Segments (HES), defined as windows whose mean token entropy exceeds the response mean. Cross-scoring reveals scorer-dependent HES coverage and overlap, while the association between HL and novelty depends on the outcome denominator and likelihood threshold. Furthermore, RLVR models have lower off-argmax ratios under self-scoring, accompanied by a higher fraction of positions with top-token probability at least 0.80. These observational findings characterize the tested RLVR models and suggest directions for improving creative generation without isolating causal training effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.