acceptodds
Under review as a conference paper at ICLR 2027

Mixture of Layers with Hybrid Attention: Parallel Thin Blocks for Sparse Transformer Compute

Abstract

We propose parallel attention processing with a new block-decomposed model, Mixture of Layers (MoL). Comparing with MoE models, the routing in MoL operates with parallel attention blocks instead of experts inside one block. MoL uses parallel thin blocks of width , connected to the full block width by learned down and up projections, and a top- routing that selects which of these blocks are processing which token. Sparse routing over multiple blocks degrade attention quality, as each block is processing fewer tokens. To address this, we introduce a hybrid attention design, with one shared block processing all input tokens using dense softmax attention (holds global context), and Gated DeltaNet linear attention in the routed blocks handling sparse token subsets. We scaled MoL to 2.08B (0.63B active) on FineWeb-Edu and found that the model outperforms an iso-active dense baseline by 0.17 PPL, while it underperforms a 1.3B dense baseline by 3.21 PPL and trains about slower. Evaluation on 8-task lm-eval-harness shows scores following PPL results, and scores model between the two dense baselines. We tested single-GPU prefill on A100, H100 SXM and H200 GPUs. At 128K context it is faster than the 1.3B dense model on A100, and up to faster at 256K tokens on H100 SXM and H200, while short-context prefill and decode remain slower.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.