acceptodds
Under review as a conference paper at ICLR 2027

Nice Fold or Hero Call: Learning Budget-Efficient Thinking under Policy-Dependent Solvability

Abstract

Large reasoning models (LRMs) improve problem solving through extended reasoning, but often misallocate test-time compute. Existing efficiency methods reduce cost by compressing reasoning traces or conditioning budget on perceived difficulty, yet read the resulting pass rate as a difficulty score, leaving the zero-return regime unmodeled. As a result, they overspend on queries beyond the model's capability while compressing hard-but-solvable ones that need deeper reasoning. In this work, we formulate adaptive reasoning as a computational investment under uncertainty, where budget follows the expected return of reasoning rather than perceived difficulty. To instantiate this principle, we propose **B**udget-**E**fficient **T**hinking (**BET**), a two-stage framework that combines behavioral cold-start with GRPO under an investment-cost-aware reward. By aligning solve-or-fold decisions with rollout-derived solvability, BET learns three behaviors: (1) *short solve*, answering easy queries concisely; (2) *nice fold*, abstaining early when continued reasoning has near-zero expected return; and (3) *hero call*, preserving sufficient compute for hard-but-solvable queries. Across seven benchmarks and three base models, BET reduces reasoning tokens by **54%** while improving up to **3.2%** accuracy, and transfers zero-shot to scientific QA and logical reasoning with comparable efficiency gains. Code is available at https://anonymous.4open.science/r/BET-CCEC/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.