acceptodds
Under review as a conference paper at ICLR 2027

BASS: A Block-wise Assessment of Sensitivity and Selective Smoothing Framework for Quantization of Branch-Conditioned Video Diffusion Transformers

Abstract

Video Diffusion Transformers (VDiTs) have substantially advanced high-resolution and controllable video generation, but their considerable computational and memory costs make practical deployment prohibitively expensive. Post-training quantization (PTQ) has proven to be an effective method for reducing memory usage and computational costs, yet existing methods primarily target VDiTs with a single denoising backbone and suffer from accuracy degradation when applied to branch-conditioned architectures with auxiliary control signals. Wan2.2-Animate-14B is a representative example: its auxiliary face and body adapters inject motion and identity conditioning signals into the main DiT, and its conditioning pathways exhibit quantization sensitivities that differ markedly from those of the backbone. Applying a uniform low-bit configuration across these components can therefore corrupt the conditioning signals, leading to motion degradation and identity drift. To address this problem, we propose BASS, a Block-wise Assessment of Sensitivity and Selective Smoothing framework for mixed-precision post-training quantization of branch-conditioned VDiTs. BASS identifies and protects sensitive layers across different branches, particularly those within the adapters, with minimal manual intervention. We validate BASS on advanced branch-conditioned VDiTs, using a mixed-precision policy combining W8A8 and W4A4, and show that it preserves motion fidelity, identity consistency, and overall visual quality under a mixed low-bit quantization policy. Furthermore, we implement efficient GPU kernels to achieve practical memory savings and inference speedups.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.