acceptodds
Under review as a conference paper at ICLR 2027

Depth-Linked Register Transformer

Abstract

Looped language models (LoopLMs) have emerged as a powerful paradigm for improving parameter efficiency while retaining strong reasoning capacity compared to fixed-depth LLMs. However, their reliance on sequential, autoregressive iterations incurs substantial computational overhead. In this work, we revisit parallelizable register embeddings for non-autoregressive latent reasoning. A key challenge of this paradigm, however, is the absence of iterative refinement and the resulting weaker sequential dependencies between latent thoughts. To address these limitations, we propose the Depth-Linked Register Transformer (DepThink), a non-autoregressive latent reasoning framework that simulates the sequential dependency and representational expressiveness of LoopLMs while preserving the efficiency of parallel inference. DepThink employs a sequence of register embeddings within a single pre-filling stage, where each register captures information from the input that is missed by its preceding registers through learnable residual connections, yielding a progressively refined latent thought. This design preserves the iterative refinement nature of the thinking steps intrinsic to looped reasoning, while inheriting the parallelism and representational capacity of register-based models, and eliminating the sequential backbone passes required by LoopLMs. Moreover, we theoretically prove that, relative to the standard transformer with registers, DepThink achieves stronger geometric directional reachability and a provable one-step descent advantage. We evaluate DepThink on reasoning-intensive text generation tasks, such as mathematical reasoning and knowledge-rich reasoning, as well as reasoning-intensive multimodal retrieval. Experimental results demonstrate that DepThink outperforms standard looped architectures and Transformers with registers across dense (2B, 8B) and sparse mixture-of-experts (30B-A3B) backbones, under both LoRA and full-parameter fine-tuning. Furthermore, wall-clock measurements reveal that DepThink can accelerate training by and inference by compared to standard LoopLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.