acceptodds
Under review as a conference paper at ICLR 2027

Not All Recurrence Is Created Equal: Depth Sensitivity Differs Between Recurrence Dimensions of a Single Checkpoint

Abstract

Recurrent language models reuse computation across inference steps, so recurrence depth appears to offer an adjustable test-time compute budget. That interpretation assumes a checkpoint remains reliable when its depth is changed after training. We test it through a controlled within-checkpoint comparison: HRM-Text-1B couples two recurrence dimensions in one set of weights, so varying each in turn holds architecture, scale, data and training schedule fixed, a comparison no cross-model study can make. Its high- and low-level dimensions turn out to respond very differently. Reducing \(L\) from three steps to two retains most trained-depth accuracy while removing 25% of recurrent block passes; increasing \(H\) from two to four reduces GSM8K accuracy from 84.0% to 10.0% and causes substantial output-format failure. Two configurations with identical total block passes differ sharply, one scores zero on every benchmark while the other does not, so total recurrent compute does not characterise the effect of altered depth. A hidden-state intervention implicates the high-level representation: additional \(H\) steps displace it from its trained endpoint while preserving its norm, and interpolating back toward that endpoint recovers most of the lost accuracy before the trained state is fully restored. We situate the result against four single-loop checkpoints from two further families, none of which shows comparable degradation over the depths we evaluate. We distil the findings into a short protocol for validating recurrence depth as an inference-time compute dial.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.