acceptodds
Under review as a conference paper at ICLR 2027

AI is Becoming an Expert Reasoner

Abstract

Across physics and chess experiments, the expertise literature shows that what differentiates experts from novices is not more search but less: experts take the necessary steps while novices employ a more circuitous trial-and-error approach. We apply this lens to large reasoning models and ask how their token efficiency has changed over time. We introduce the minimal human derivation, the shortest published human solution to a problem, as a problem-specific benchmark against which excess reasoning can be measured across problems and time. On 40 problems from the 2026 AIME and HMMT competitions, evaluated across 15 frontier models with repeated sampling, mean tokens per correct solution fell 8.5x across OpenAI models (o1 to GPT-6-Astra) and 4.8x across Anthropic releases (Opus 4.5 to Fable 5.1), while accuracy rose in both families. Excess reasoning above the human benchmark falls 22.3-32.6% for OpenAI and 31.7-45.4% for Anthropic per quarter. Extrapolating within family, average frontier traces come within 10% of the human benchmark by late 2027 to late 2028. Matched pairs of open-weight models suggest efficiency gains come from both pre-training scale and post-training. We discuss implications for forecasting, the token economics of AI, and chain-of-thought monitorability as reasoning approaches the length of the shortest human derivation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.