Reflect Less, Reason Better: Entropy-Aware Reflection Rewards for Efficient Large Reasoning Models
Abstract
Large reasoning models make significant progress in solving complex problems, but their reasoning processes are still often affected by overthinking, which leads to high reasoning latency and computational costs without corresponding improvements in accuracy. Existing methods usually improve reasoning efficiency by directly penalizing reasoning length or the number of reflection words through reinforcement learning. However, these methods often ignore differences between problem difficulty and reflective reasoning structures, which may cause models to frequently switch reasoning paths and thus produce fragmented reflection. To address these issues, this paper proposes a reflection reward method that integrates answer entropy and difficulty awareness, based on analyses of the number of reflection words, continuous reasoning spans, and answer distributions. This method focuses on constraining inefficient reflection behaviors in two types of scenarios: high accuracy with low entropy, and low accuracy with high entropy. It penalizes fragmented and cyclic reflection, and favors reasoning trajectories with low-frequency reflection, long continuous reasoning spans, and continuous stability. Extensive experimental results show that, without reducing model performance, the proposed method effectively reduces token consumption in chain-of-thought reasoning processes, providing a more fine-grained and effective optimization approach for balancing efficiency and accuracy in large language model reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.