UniQ: Modality-Aware Post-Training Quantization for Unified Multimodal Models
Abstract
Unified Multimodal Models (UMMs) integrate visual understanding and generation within a single Transformer backbone, but their large parameter count creates a substantial deployment bottleneck. Post-training quantization (PTQ) is a natural solution to this problem. However, directly applying quantization methods developed for language models, vision-language models, or diffusion models proves suboptimal because these methods do not match the token routing and dual-task structure of UMMs. We find that in UMMs token types and task pathways exhibit distinct statistical behavior, making calibration, precision allocation, and activation quantization unreliable. To address these challenges, we present UniQ, a modality-aware PTQ framework for unified multimodal models. UniQ introduces three innovative designs: dual-modality calibration, pathway-adaptive mixed precision, and token-aware sparse channel rotation. Experiments on BAGEL show that UniQ achieves W3A4 quantization with negligible quality loss and a 2.9 speedup over full precision. Our code and model will be available to the public.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.