acceptodds
Under review as a conference paper at ICLR 2027

Dynamic FFNs Improve Representation Learning in Transformer Pretraining

Abstract

Modular architectures allow different inputs to invoke different computations, but existing approaches often couple this flexibility with increased model capacity. We ask how far input-dependent modularity can take us when parameter and compute budgets are held fixed. We introduce ynamic eedorward etworks (DFFNs), which reorganize a standard Transformer feedforward block into a shared static FFN and a set of routed low-rank Submodules (SMs). A learned router selects an SM for each token, providing input-dependent computation without increasing the parameter or compute budget of the dense FFN. Across multiple model scales, DFFN consistently improves language modeling and downstream performance over matched dense Transformers. Controlled experiments further show that these gains depend on learned input-dependent routing rather than simply the presence of the submodules. Together, our results show that modularity can improve Transformer pretraining without relying on increased model capacity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.