acceptodds
Under review as a conference paper at ICLR 2027

TurboQuantConv: Channel-Isolated Incoherence Processing for 3D Convolutions in Video Generative AI

Abstract

Three-dimensional Convolutional Variational Autoencoders (3D-VAEs) serve as the fundamental compression backbone of modern video generative models, yet their deep 3D convolutions (Conv3D) impose prohibitive computational and memory burdens. Pushing post-training quantization (PTQ) down to 4-bit weights and activations (W4A4) typically causes substantial quality degradation due to non-stationary spatiotemporal activation bursts. In this paper, we propose TurboQuantConv, a channel-isolated incoherence processing and quantization framework for deep multi-dimensional convolutions. By harmonizing outlier suppression with convolutional receptive field geometry, TurboQuantConv effectively suppresses spatiotemporal bursts and achieves high-fidelity quantization without requiring fine-tuning or retraining. Under W4A4 compression, TurboQuantConv outperforms prior state-of-the-art PTQ methods by  4 to 8 dB PSNR across 3D-VAE benchmarks. In downstream 81-frame video generation with the 5B Wan 2.2 Diffusion Transformer, TurboQuantConv achieves the lowest FVD while closely preserving unquantized visual dynamics, avoiding the severe motion freeze and chromatic collapse from which existing PTQ methods suffer. These results establish a practical foundation for efficient edge deployment of video generative models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.