LIAR: Length-Interval-Based Reward Shaping for Reliable and Accurate Chain-of-Thought Reasoning
Abstract
Large Reasoning Models (LRMs) have achieved remarkable success in complex problem solving through Reinforcement Learning with Verifiable Rewards (RLVR), which encourages them to generate chain-of-thought (CoT) reasoning traces. However, CoT length critically influences performance: excessively long traces increase inference cost and accumulate errors, while overly short ones break logical coherence and degrade accuracy. Although prior work suggests that the optimal CoT length varies with problem difficulty, directly targeting a single optimal length during policy learning remains impractical. In this paper, we propose the (ORI), a statistically grounded interval of near-optimal reasoning lengths that contains the theoretical optimum and maintains high reliability across its range. Building on ORI, we introduce ength-nterval-bsed eward shaping (), which applies a piecewise-linear length penalty to correct trajectories outside a data-driven ORI estimate, thereby guiding reasoning toward a reliable and high-accuracy length regime. Extensive experiments across four base models show that steering generation into a length interval—rather than compressing toward the shortest one—yields higher accuracy than prior CoT compression methods: in our primary DeepSeek-R1-Distill-Qwen-1.5B setting, LIAR improves over the strongest compression baseline LASER-D by \textbf{+1.7\\%} on average, with the largest gains on the hardest benchmarks (\textbf{+2.7\\%} on AIME 2026), where long reasoning is needed to stay reliable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.