acceptodds
Under review as a conference paper at ICLR 2027

DIET: Deletion-response Expert Trimming for Video Diffusion Transformers

Abstract

Video diffusion transformers (DiTs) increasingly rely on mixture-of-experts (MoE) architectures, where only a sparse subset of a large expert bank is activated per token. While dynamic sparsity reduces active compute, it leaves the full expert storage footprint intact. Furthermore, conventional one-shot pruning criteria rely on static activation or routing statistics, failing to capture the layer-level re-routing behavior triggered after expert deletion. To make this post-deletion behavior observable without prohibitive cost, we record all expert outputs and router states across matched conditional and unconditional tokens in a single all-expert calibration pass. Consequently, replaying any single-expert deletion and its grouped re-routing reduces to tensor arithmetic on cached states, requiring zero additional model forward passes. Using these deletion responses, we formulate retained-expert selection as an Overall Diversity Loss (ODL): each pruned expert is matched with its nearest retained expert and the summed nearest-neighbor cosine distances define a set-level coverage loss under which deleted experts closer to retained ones cost less, directly prioritizing functional coverage. We solve this optimization in two stages: an intra-layer local search combining greedy initialization, single-swap refinement, and simulated annealing to select retained set under a fixed retention ratio, and an inter-layer regression-guided search that optimizes expert-count allocation across layers. Evaluated on LingBot-Video 30B-A3B, pruning 50% of the experts (6,144 → 3,072) without fine-tuning reduces the checkpoint footprint from 57 GB to 30 GB, enabling single-card deployment on a 48 GB GPU while generation quality remains essentially intact: we observe the official VBench Total rise from 0.7941 to 0.8115 across a fixed 284-case protocol. Across all other tested retention budgets, DIET consistently outperforms competitive baselines adapted from large language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.