Don't Waste Failures: Leveraging Failure Length Signals for Discriminative Thresholds in Efficient Reasoning
Abstract
Large reasoning models (LRMs) improve problem-solving accuracy through extended chains of thought, but can incur substantial cost from overly long reasoning. Existing reward-based approaches to reasoning efficiency often set length penalties using overall or success-only length statistics, failing to carefully differentiate how reasoning lengths differ between correct and incorrect rollouts. We show that leveraging this difference informs the design of length control: When failures are shorter than correct rollouts, we retain a higher penalty threshold, whereas longer failures motivate an earlier penalty. Based on this insight, we introduce OASIS (Outcome-Aware threshold Selection with Integrated Signals). OASIS estimates the reasoning-length threshold from the length distributions of correct and incorrect rollouts, determining where the length penalty begins. We further introduce a soft directional estimator to capture short and long failures separately, increasing the threshold for the former and decreasing it for the latter. Because a higher threshold leaves more reasoning unpenalized and weakens the effective length penalty, OASIS rescales the penalty gain to compensate for this weakening of the compression pressure while leaving the selected penalty-free region unchanged. Integrated into a length-aware GRPO reward, OASIS effectively reduces reasoning length relative to the base model across four base models evaluated on MATH500, AIME 24–26, and AMC23, while largely maintaining accuracy and yielding an average improvement of 1.7% points. Notably, on DeepSeek-R1-Distill-Qwen-1.5B, OASIS improves average accuracy by 4.5% points while reducing average reasoning length by 78.2%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.