acceptodds
Under review as a conference paper at ICLR 2027

Colinearity Decay: Training Quantization-Friendly Transformers via Decoupled Weight-Space Regularization

Abstract

Low-bit quantization enables efficient Transformer deployment, yet activation outliers complicate fully quantized inference. Quantization methods act after or during training: post-training methods adapt pretrained models, while many training-time approaches address outliers through architectural changes, or by regularizing activations. Yet minimizing activation magnitude is incomplete: excessive suppression can degrade full-precision accuracy without consistently improving quantized models. We argue that quantization-friendly training should target harmful structural artifacts that cause extreme activations. To this end, we introduce Colinearity Decay (CD), a weight-space regularizer for matrix pairs within Transformer blocks. CD penalizes detrimental cross-matrix alignment and mitigates extreme activations. The decoupled update leaves architecture and task loss unchanged, adding little training overhead and no inference-time cost. Across ImageNet-1K pre-training, object detection, downstream fine-tuning, and small-scale LLM pre-training, CD improves quantized task performance across algorithms and bit widths while preserving unquantized performance. These results support decoupled weight-space regularization for training quantization-friendly Transformers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.