PebbleVideo: On-Device Video Generation via Extreme Model Compression
Abstract
Diffusion Transformers (DiTs) have substantially advanced video generation quality, yet their quadratic bidirectional attention and billion-scale parameter incur significant memory and storage overhead, hindering deployment on resource-constrained devices. Existing approaches improve efficiency through attention redesign or streaming generation, but their restricted global spatiotemporal interactions limit modeling capacity. In contrast, we explore an alternative direction, model compression, and introduce **PebbleVideo**, which aggressively reduces the parameter footprint through structured compression. The core of PebbleVideo is to evaluate compression sensitivity in functional space, preserving generation-critical functions rather than relying on indirect proxies. For the dominant DiT backbone, we measure structural importance through the perturbation induced on the diffusion vector field, directly exposing each module's influence on generation dynamics. For heterogeneous condition encoders, cross-scale representation consistency determines whether compression proceeds through compact-model substitution or in-place structural pruning. Combined with DMD-based distillation, PebbleVideo achieves 512320 video generation at 28 FPS with only 7.45 GB peak memory. Further incorporating model quantization, we deploy the quantized PebbleVideo on a Samsung mobile NPU, achieving a VBench score of 87.46.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.