acceptodds
Under review as a conference paper at ICLR 2027

T-LoopFormer: Token-Level Elastic-Depth Looped Transformers with Dynamic Routing

Abstract

Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency without sacrificing representational power. Besides, looped Transformers perform inference directly in the latent space (latent reasoning) to reduce the number of tokens consumed during inference, thereby achieving improved sample efficiency. However, these models typically apply a fixed recursion depth uniformly to every token, leading to suboptimal compute allocation and leaving significant efficiency gains on the table. In this work, we propose token-level elastic-depth looped transformers: T-LoopFormer. T-LoopFormer uses dynamic token-choice routing, which enables each token to adaptively determine its own number of loop iterations based on its hidden state, which could improve the token generation accuracy. Moreover, we further introduce recursion-wise KV cache, which maintains an independent key-value cache for each recursion loop of active tokens; this design ensures that tokens at different depths only attend to their corresponding cached states, effectively enabling faster autoregressive decoding. Extensive experiments show that T-LoopFormer achieves robust performance on a wide range of language modeling and zero-shot reasoning tasks, and our model could reach the lowest decoding latency, which validates the effectiveness of token-choice router and recursion-wise KV cache.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.