ARLoop: Strong Fully-Open Looped Transformers with Attention Residuals
Abstract
Looped transformers increase effective depth by repeatedly applying a shared set of transformer blocks, offering a parameter-efficient way to scale computation. However, standard residual connections aggregate information across recurrent depth into a single stream, providing limited control over which intermediate representations from preceding layers should be reused. In this paper, we introduce ARLOOP, a looped transformer architecture based on attention residuals, where the input of each layer is selectively aggregated through attention over model depth. Distinct from standard transformers, we find that looped transformers do not require access to the full residual history: sliding-window depth attention matches or outperforms full-history attention at larger recurrence depths. Across controlled experiments with multiple parameter scales (140M and 370M) and recurrence depths (–), ARLOOP consistently outperforms prior residual designs for looped transformers. More importantly, we further scale ARLOOP to 0.6B and 1.3B parameters under a realistic LLM pretraining recipe, demonstrating that the gains transfer to realistic pretraining, with particularly strong improvements on math and coding tasks. Our 1.3B model, pretrained on only 0.5T tokens with three recurrences, outperforms the fully-open looped LLM Huginn-3.5B across all evaluated tasks and is competitive with similar-sized non-looped models such as Qwen3.5-2B despite being trained on substantially less data. ARLOOP will be fully open: we will release its code, checkpoints, and training recipe.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.