MorphoQuant: Morphology-Driven Quantization for Omni-modal Large Language Models
Abstract
Conventional Post-Training Quantization (PTQ) methods struggle with 4-bit Omni-modal Large Language Models (OLLMs) due to the extreme distribution heterogeneity and disparate outlier patterns across modalities. To address this, we propose MorphoQuant, a morphology-directed PTQ framework engineered to preserve cross-modal morphology and mitigate outlier loss. Specifically, we introduce Morphology-Driven Sparse Compensation (MDSC), which identifies high-dispersion activation channels and retains large activation residuals into a channel-wise sparse computational branch. Complementing this, we propose Morphology-Preserving Optimization (MPO), which optimizes the activation clipping boundary using a reconstruction objective with a cosine-similarity term; the boundary also determines which residuals enter the compensation branch. We evaluate the method on Qwen2.5-Omni-3B and 7B using vision-language, video-language, and audio-language benchmarks, and on InternVL2.5-8B using vision-language benchmarks. Under the evaluated W4A4 setting, MorphoQuant maintains non-collapsed task performance where the reproduced W4A4 multimodal baselines frequently fail. Alternative activation formats further improve results in some settings. These findings suggest that fine-grained residual selection coupled with clipping-boundary optimization is a promising approach to low-bit quantization of omni-modal models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.