acceptodds
Under review as a conference paper at ICLR 2027

When Is More Reasoning Worth It?

Abstract

The next generation of reasoning models will not be the ones that think longest, but the ones that know when to think. Adapting reasoning effort to each question is a central open problem, and solving it would move the frontier between cost and accuracy. This work is a step in this direction. We build a controlled evaluation of effort selection and release its data: repeated responses from open models at every effort setting on the same questions. We first show that adapting effort to the question can save nearly half of the generated tokens with almost no loss in accuracy. We then compare two practical strategies, predicting from the question text and checking agreement among cheap-model answers, with every model call charged. We propose question-level cost forecasts that improve an existing selector without extra generation. Finally, we identify where further gains lie: knowing where extra reasoning helps matters beyond knowing where it is expensive, but this benefit is model specific and current forecasts capture little of it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.