Boltzmann-MoE: Energy-based mixture-of-experts with native routing
Abstract
In spite of the dominance of transformers in advanced AI models, our theoretical understanding of their inner workings from a statistical modeling perspective is still incomplete. In this work we show that transformers may be related to a very fundamental class of statistical models. Taking a first principles approach, we introduce the "Boltzmann-MoE", an energy-based model to learn both the overall data distribution, and which adapts itself to the context tokens. Boltzmann-MoE is a generalization of the Energy Transformer (ET) – a variant of looped transformers with some weight-tying – and naturally includes a mixture-of-experts with a native router. The Boltzmann-MoE energy is surprisingly simple and intuitive, consisting of two "free energies" (negative log-partition functions): the ET attention energy, which is a KDE on the context tokens, and the MoE energy is the log sum of Boltzmann weights measuring alignment of the token with memorized patterns. The forward pass of Boltzmann-MoE is a preconditioned gradient update, yielding softmax attention, softmax gating of the experts, and experts that combine a two-layer MLP with a gated linear unit. With a preconditioner split into dissipative and rotational parts, this update is a port-Hamiltonian flow whose learned rotational part adds deterministic exploration of the energy landscape while keeping the forward pass end-to-end differentiable. Our model is an alternative to standard transformers and MoE grounded in statistical physics. At M–M active parameters on Nemotron-CC, hybrids with a standard transformer encoder followed by Boltzmann-MoE blocks, which clearly outperform pure energy models, trail an iso-active Switch-MoE by – points of average zero-shot accuracy and need about twice its tokens to reach the same loss. In deep models, distinct energy blocks learn update fields – more aligned than those of standard transformer layers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.