Scaling Laws for Latent Reasoning: How Test-Time Compute Tracks Circuit Depth
Abstract
A recurrent-depth language model spends test-time compute by iterating a weight-shared block times, reasoning in latent space rather than emitting tokens. We show that this depth axis obeys a mechanistic scaling law: latent accuracy is the cumulative distribution function of a task's circuit depth, discounted by a per-step reliability that grows with model size . The law makes three falsifiable predictions: (i) accuracy saturates in , (ii) the iterations required grow linearly with reasoning depth, and (iii) the value of additional computation is proportional to the serial depth it advances. We verify these predictions on a compute-optimal ladder trained from scratch (M–M non-embedding parameters) across three reasoning domains. Fitting the law on smaller models predicts the entire depth-accuracy curve of the held-out largest model (), with the deeper of two depths requiring more iterations in of untied comparisons. The released Huginn-B follows the same saturating law beyond our ladder on GSM8K, ProsQA, and SVAMP (), while GSM8K accuracy decays geometrically with gold-solution depth (). Continuous-thought positions provide substantially less accuracy gain per unit of compute than recurrent depth, and Coconut does not out-scale CoT on public ProsQA. Optimizing the law yields an inference-time analogue of the Chinchilla rule for allocating compute between model size and reasoning depth. Across these settings, test-time compute scales when it advances the serial depth required by the task, establishing circuit depth as the relevant unit of latent computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.