acceptodds
Under review as a conference paper at ICLR 2027

T²MLR: Transformer with Temporal Middle-Layer Recurrence

Abstract

Transformer reasoning is limited by autoregressive decoding, which repeatedly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time. We introduce Transformers with Temporal Middle-Layer Recurrence (T²MLR), a Transformer-based latent reasoning architecture that fuses a cached middle-layer representation from the previous token directly into an earlier layer of the current token position, enabling abstract intermediate computation to persist across decoding steps with little inference overhead. Across natural-language pretraining and multi-hop reasoning finetuning, middle-layer T²MLR variants outperform data- and parameter-matched Transformer baselines in every setting we test, with gains that persist when scaling to 1B parameters or 50B pretraining tokens. Moreover, applying recurrence to only a localized middle-layer block (as little as 20% of the network) often outperforms full-layer recurrence. Importantly, T²MLR does not require pretraining from scratch: retrofitting the recurrent pathway into an existing pretrained 1.7B Transformer and briefly finetuning substantially improves math reasoning, lowering the barrier to practical adoption. These results suggest that effective latent reasoning in Transformers does not require recurrence through all layers, as in prior work, but can instead emerge more strongly from targeted middle-layer recurrence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.