acceptodds
Under review as a conference paper at ICLR 2027

ESEA: Error-Aware Sparse Expert Allocation for Efficient Diffusion MoE Inference

Abstract

Mixture-of-Experts (MoE) has emerged as a promising paradigm for scaling diffusion models, yet activating a fixed number of experts for all tokens leaves computational redundancy. Existing expert skipping methods mainly target LLMs and overlook the iterative denoising nature of diffusion models, leading to degraded generation quality when directly applied to diffusion moe models. We analyze the underlying error mechanisms and identify three insights: (i) approximation error depends on expert similarity; (ii) local errors exhibit timestep–layer heterogeneity and cross-prompt consistency; (iii) the impacts of local errors on global error vary across timesteps. Based on these insights, we propose ESEA (Error-aware Sparse Expert Allocation), a training-free method for diffusion MoE inference. ESEA first adjusts routing weights based on expert similarity to reduce approximation error, then weights calibrated 3D local errors tensor to reflect their timestep-dependent impacts on global error. Finally, dynamic programming minimizes the weighted local error, jointly allocating experts across timesteps and layers under a fixed total expert budget. The routing weight adjustment strategy and expert allocation are determined offline, introducing negligible overhead during inference. Extensive experiments on HunyuanImage-3.0-Instruct, its distilled variant, and LingBot-Video demonstrate that ESEA achieves the best performance among the compared methods under the same expert budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.