RecipeMoE: Learning Reusable Block Recipes for Full-Participation Mixture-of-Experts
Abstract
Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. Participation is how many experts are involved in producing a token's output, execution is how many are evaluated on the token, and materialization is how many expert-sized parameter sets must be built. Sparse routing keeps execution and materialization low, but shrinks participation: for each token, only a few experts participate. Dense output-mixing restores full participation, but its execution grows with the number of experts. Parameter-merging keeps execution at one expert, but its materialization grows with the number of routing decisions. We propose RecipeMoE, a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution. Its cacheable blocks come from a small learned codebook, one per entry. A lightweight hypernetwork generates block-specific composition recipes that merge each internal layer's expert pool into a composed expert. Participation is full, because every composed expert is synthesized from the entire pool. Execution stays sparse, because a router sends each token to only a few blocks. Materialization is bounded, because the codebook, not the input, fixes how many blocks exist. Dual-Path Residual Gating (DPRG) enables nonlinear interactions between two paths composed from the same expert pool using independent recipes. Experiments on image classification show consistent gains over representative sparse and dense MoE baselines. Additional experiments on language modeling and sequential recommendation validate its generalization beyond vision. RecipeMoE is fully deployed in a large-scale industrial generative recommendation system, serving hundreds of millions of users under a 60 ms latency budget, with a 2.4% relative UVCTR gain in online A/B testing. Code is available at https://anonymous.4open.science/r/RecipeMoE-717C254/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.