acceptodds
Under review as a conference paper at ICLR 2027

Tool Use Reduces Depth-Induced Collapse in OOD Reasoning

Abstract

Humans can apply ideas learned in one context to substantially different situations. We call this process of searching for and constructing novel recombinations of learned relationships to solve new problems out-of-distribution (OOD) reasoning. The capacity for large language models (LLMs) to support OOD reasoning underpins proposals for generally intelligent systems. However, this property is challenging to measure because most problems admit many decompositions, some involving shallow subproblems and others involving subproblems that may have been memorized from the training data. This makes it difficult to determine how much compositional reasoning a model must actually perform. Uncertainty about training distributions, how to measure a datapoint's distance from a training distribution, and the exponential number of ways to decompose most tasks make it intractable to robustly measure any modern LLM's OOD-reasoning capacity on standard natural-language tasks. In this work, we introduce a benchmark in which a model progressively solves a Boolean circuit over from data. This benchmark is both minimal, isolating OOD reasoning from these confounding factors, and general, as any computable function can be represented as a sufficiently large polynomial. We find that standalone models’ next-step accuracy collapses as depth grows. In contrast, tool use through the synthesis and execution of code prevents this collapse in both small and frontier LLMs. These results indicate that synthesizing tools play a crucial role in supporting OOD reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.