RazorReward-RL: Capability-Aware Efficiency Optimization for LLM Reasoning via Razor Reward
Abstract
Chain-of-Thought (CoT) prompting is a de facto standard for unlocking strong reasoning performance in large language models (LLMs), but its step-by-step generation often introduces redundant computation in cost- and latency-sensitive production deployments. Existing reinforcement learning (RL) approaches for CoT efficiency rely on length-based rewards to mitigate redundant steps, but overlook the model's intrinsic query-specific reasoning capacity. This leads to two key limitations: for Beyond-capability queries that exceed the model's reasoning capacity, the reward mechanism encourages longer but futile reasoning, wasting massive token budgets; for Near-capability queries lying at the edge of the model's capacity, smoothed length-based rewards fail to provide clear discriminative signals to identify optimal stopping points, causing the model to either stop early with insufficient reasoning or drift away from the correct answer after redundant reasoning. To address these issues, we propose RazorReward-RL, a capability-aware RL framework for efficient CoT reasoning. As its name suggests, the core RazorReward function acts as a razor for reasoning chains. By segmenting CoT trajectories at natural reasoning step boundaries to construct semantically coherent prefixes, it sharply amplifies the reward gap in favor of the shortest sufficient prefix, i.e., the shortest prefix that yields a correct answer, imposes steep penalties for under- and over-reasoning on all within-capability and near-capability queries, while imposing heavy penalties for generation on beyond-capability queries to block futile reasoning. Experiments across multiple domains and model scales show that RazorReward-RL significantly reduces token consumption while improving accuracy, achieving a better accuracy-efficiency trade-off than prior methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.