QuIL: Quantifying Quantization-induced Information Loss in Neural Networks
Abstract
Floating-point bit-string representations of parameters are ubiquitous building blocks of neural networks, and their quantization into lower precision formats is increasingly used to reduce their computational requirements. The degradation that quantization induces in the model is, however, not well studied, and its detection relies on task-specific heuristics. We propose QuIL: An information-theoretic approach for measuring quantization-induced information loss. By modelling the probability distribution of the underlying floating-point bit-strings of the parameters with a probabilistic circuit, we can measure information loss as the change in entropy across quantization levels. Experiments on neural-network regression, vision transformers, and GPT models show that QuIL, computed from the parameter posterior alone, identifies the bit-width at which quantization begins to degrade accuracy, calibration, and predictive density, offering a principled basis for selecting parameter precision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.