acceptodds
Under review as a conference paper at ICLR 2027

How Deep is Your Thinking? Evaluating the Depth and Efficiency of Long Horizon Reasoning

Abstract

Large language models (LLMs) with chain-of-thought reasoning capabilities have driven a step change in the classes of problems that are solvable with foundation models. In particular, the vanilla Transformer architecture is limited in the number of sequential computations it can perform within a fixed number of layers, whereas chain-of-thought reasoning greatly expands this capability in a manner that adapts to the problem complexity. However, the precise extent of this increase in capability, as well as the efficiency with which it is achieved, has not yet been systematically explored. To address this gap, we propose a new evaluation suite, DepthBench. DepthBench aims to measure the effective reasoning depth of models i.e. the number of sequential computations they are reliably capable of performing during inference. We focus on the domains of function composition for arithmetic and state tracking for planning, where we have direct control over problem complexity in terms of the required sequential computations necessary for success, allowing us to study how LLMs with and without chain-of-thought reasoning perform as a function of this complexity. We find that prominent open-source LLMs exhibit increases of one to three orders of magnitude in the achievable sequential computation depth, depending on the domain. We also find that models in a medium parameter range are the most efficient in their use of layers to increase depth, both with and without chain-of-thought. DepthBench enables the systematic evaluation of the effective depth of LLMs that employ adaptive computation and helps foster well-grounded innovation on the limits and efficiency of these models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.