acceptodds
Under review as a conference paper at ICLR 2027

From Uncertainty to Consequence:Learning from Semantic Forks in LLM Reasoning

Abstract

Reinforcement learning with verifiable rewards (RLVR) has emerged as an important approach for improving the reasoning capabilities of large language models (LLMs). Recent studies have further explored token-entropy patterns and their value for training, viewing high-entropy tokens—where the model hesitates among possible continuations—as potential branching points in reasoning. However, the uncertainty of an individual token does not reveal the specific reasoning path selected by the model. Different reasoning paths and their divergences become observable only when the analysis is extended to the level of semantic segments. Based on this insight, we propose SemCredit, a fine-grained credit assignment method for natural-language reasoning. SemCredit organizes multiple completed rollouts for the same problem into a semantic prefix trie. It identifies distinct reasoning paths through shared semantic prefixes and their subsequent branches, and estimates local advantages from terminal outcomes aggregated at the corresponding nodes. By combining trajectory-level rewards with semantic branch information, SemCredit enables more targeted credit assignment without requiring step-level supervision. We further investigate the positional effects of high-entropy tokens during model reasoning. Across four Qwen models of different scales and six mathematical reasoning benchmarks, SemCredit achieves the highest average performance among all methods evaluated under the controlled setting. On Qwen3-4B, it outperforms the strongest controlled baselines by 9.07 percentage points on AMC 2023 and 8.75 percentage points on AIME 2024. In medical reasoning experiments, SemCredit also achieves the best performance across all three evaluated benchmarks. Overall, these results demonstrate that jointly modeling the positions of high-entropy decisions and semantic-level reasoning paths provides a more structured understanding of RLVR, enabling the identification of critical reasoning segments and the decisive tokens within them for more effective optimization of LLM reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.