acceptodds
Under review as a conference paper at ICLR 2027

Can RLVR Models Think More Creatively? A Study of Overconfidence and High Entropy Segment

Abstract

Recent advances in Large Language Models (LLMs) are often attributed to improved reasoning abilities through Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR). Although RLVR models perform strongly on mathematical benchmarks, their ability to generate creative solutions remains an open question. In this work, we compare RLVR models and math-pretrained models from the DeepSeek and Qwen families on mathematical problem solving with supplied reference solutions. RLVR models yield fewer responses judged both correct and novel, while their self-scoring high lower-tail token likelihood (HL) rates exceed those of the Math models by 15.85 and 74.70 percentage points, respectively. HL operationalizes token-level Overconfidence (OC) in this study, rather than calibration error. Motivated by the hypothesis that SFT and policy optimization reshape token probabilities, we analyze High Entropy Segments (HES), defined as windows whose mean token entropy exceeds the response mean. Cross-scoring reveals scorer-dependent HES coverage and overlap, while the association between HL and novelty depends on the outcome denominator and likelihood threshold. Furthermore, RLVR models have lower off-argmax ratios under self-scoring, accompanied by a higher fraction of positions with top-token probability at least 0.80. These observational findings characterize the tested RLVR models and suggest directions for improving creative generation without isolating causal training effects.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.