acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Overthinking: Effects of Run-to-Run Instability in Large Reasoning Models

Abstract

Large Reasoning Models solve complex tasks by generating intermediate steps before outputting a final answer. However, they often exhibit a failure mode known as *overthinking*, intuitively defined as generating excessively long reasoning chains for a problem the model could otherwise solve concisely. The literature traditionally diagnoses this by checking whether a single reasoning trace exceeds a fixed token budget. This single-trajectory view ignores the complexity of stochastic sampling. To overcome this limitation, we generate 100 independent runs per problem, revealing high variance in trace length. We reframe overthinking as *instability* and introduce a difficulty-aware, threshold-free framework to measure it. We define a model-relative measure of *empirical difficulty*, a new overthinking-quantifying metric, and we use a variance decomposition to isolate within-problem instability from between-problem difficulty. Applying this framework across three datasets (MATH500, GSM8K, DUMB500 Math) and seven models (Gemma4 and Qwen3, from B to B parameters), we find that models think more than they *could* and *should* in comparison to their ability. We show that runs longer than usual for their problem are systematically less likely to be correct, and for some models the relationship between accuracy and relative length takes an inverted-U shape. This instability does not scale monotonically with parameter count, and it concentrates on easy-to-medium problems before collapsing on the hardest ones. We suggest that future evaluation protocols report this run-to-run instability alongside mean length and accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.