Single-Layer, Two Axes: Communication and Refinement in Looped Transformers
Abstract
Looped Transformers increase computation by reusing parameters and outperform standard Transformers of the same size, but their architectures often change two things together: how a token is processed and which intermediate representations it can read. This coupling obscures the source of their gains. To disentangle the two, we introduce a Single-Layer framework that separates refinement, the choice of feed-forward operators, from communication, access to historical token–depth states. Unfolding recurrent depth into an addressable sequence of latent states lets one shared attention module support controlled changes to either axis. In 274M-parameter looped MoE language models with matched FFN compute, cross-depth access lowers held-out perplexity by 2.4–8.1% relative to same-depth access under all four refinement schedules, yet depth-specific expert banks still outperform a shared bank by 11.8–15.5%. Better communication and more effective parameter reuse are therefore distinct design questions, and they interact: the benefit of start injection depends on the available historical context. Varying the retained depths for recent and distant tokens shows that the preferred depth depends on token age and refinement, and that retaining all tested depths is worse than keeping the best one. These observations motivate a Single-Layer Transformer combining routed expert refinement with recent working memory and selective long-term memory. A compact 20M-parameter comparison at matched total capacity confirms the cross-depth advantage under both bank assignments, with a smaller gap between shared and depth-specific refinement. In the shared-bank model, W = 8 lowers perplexity by 7.9% relative to same-depth access and by 0.37% relative to full cross-depth access, while storing 84.8% fewer historical states.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.