CROWN: Conversion to Routed Experts with Output-Aware Width Narrowing for Audio-Video Diffusion
Abstract
Feed-forward networks (FFNs) account for a substantial fraction of the param eters and per-step computation in video diffusion transformers, yet their intra expert redundancy remains underexplored in multimodal token-routing settings. We start from DaVinci-MagiHuman (DaVinci), a pretrained single-stream audio video diffusion transformer, and sparse-upcycle the shared MLPs in its 32 mid dle blocks into token-level Top-1 Sparse-MoE layers, while leaving the modality specific early and late blocks unchanged. This conversion restructures the shared FFNs into routed expert subspaces without reducing the active intermediate width. Building on TENP, we propose Elastic Expert Neuron Pruning (EENP), a modality-aware method for structured intra-expert pruning in routed multimodal DiTs. EENP estimates intermediate-neuron importance separately from video, audio, and text tokens, normalizes the aggregated statistics within each modality, and fuses them into a modality-aware score with weights 3:1:1 for video, audio, and text. By aggregating statistics within each modality before fusion, the scoring procedure prevents modalities with more tokens from dominating the pruning decision. Low-importance neurons are then structurally removed, with the retained neurons stored as a nested Fast prefix of a wider Quality width. Under a 60% nominal pruning ratio,EENP retains 92.4% of the unpruned SparseMoE VBench- 2.0 score and reduces DiT denoising latency by 22.0% relative to the unpruned SparseMoE checkpoint.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.