Mambura: a Mamba-Transformer Hybrid for Unified Multimodal Understanding and Generation
Abstract
Unified models for image understanding and generation face competing task objectives and the quadratic cost of long-context attention. We introduce Mambura, a state-space model (SSM) architecture that, to our knowledge, is the first Mamba–Transformer hybrid to unify these tasks. Its core Modality-aware Mixture-of-Mamba-Experts (M³) layer combines modality-specific expert projections with a shared state-space core, selecting one expert per token within its modality's expert pool. A small number of attention layers provide global token interactions. Its principal advantage is inference efficiency. Against Qwen3.5-2B, the only baseline that still runs at a 1M-token context, Mambura prefills 2.3× faster at 128K and 15.8× faster at 1M, decodes 4.6× faster at 128K and 1.7× faster at 1M, and holds 26% lower peak memory at 1M. The full-attention unified models exhaust the device much earlier. Activating only 0.39B of its 1.7B parameters per token, Mambura nonetheless outperforms dense Transformer baselines that activate seven to eight billion: it exceeds Chameleon, UniToken and Emu3 both on GenEval, where it reaches 0.70, and on MMMU, and it outperforms Show-o on three of their four shared understanding benchmarks. These results demonstrate the potential of SSM-based hybrid architectures for efficient unified multimodal modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.