Mobius Learning: Cyclic Depth Folding in Transformers
Abstract
Transformer-based language models organize computation along an ordered depth axis, where shallow and deep blocks often develop distinct representational roles. We introduce , a training architecture that allows the same block parameters to serve multiple depth roles. Through cyclic depth folding, different data streams traverse shared block groups in cyclically shifted orders, jointly training each group in shallow and deep positions, a phenomenon we call depth-role superposition. In mixture-of-experts (MoE) language-model pretraining, combines effectively with looping, which reuses parameters through repeated passes over the same block sequence. At eight loop passes, achieves higher five-task zero-shot accuracy than standard looping across all four evaluated model sizes. It also supports effective data reuse: under eightfold repetition, raises five-task zero-shot accuracy from 46.13% to 48.29% relative to standard looping for the largest evaluated model. These findings establish depth-role superposition as a viable training mechanism that combines with looping and improves downstream performance under repeated-data training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.