The Butterfly Effect in Large Language Model Quantization
Abstract
Quantization has been widely used to reduce the memory footprint and inference latency of large language models (LLMs). Low-bit quantizers have typically been evaluated using average metrics such as perplexity and task accuracy. Nevertheless, these average metrics do not reveal failures on individual inputs, such as correct answers becoming wrong after quantization. Focusing on tail reliability, we propose a unified source-to-decision analysis, which measures rare output failures and traces each input's quantization errors from individual modules through downstream amplification to changes in output distributions, token decisions, and autoregressive generation. Our analysis reveals the quantization butterfly effect: High-gain computation paths amplify small rounding errors on particular inputs into large output shifts and token changes. These amplified errors form an upper tail hidden by average metrics, and an early token change redirects the autoregressive continuation. Our experiments demonstrate the quantization butterfly effect on Qwen2.5-3B, where 4-bit weight-only GPTQ increases mean perplexity by only 7.3%, yet the 99th-percentile (p99) per-token KL divergence is 8.9 times the mean, 14.8% of top-1 predictions flip, and almost every greedy continuation leaves the full-precision trajectory. The same pattern recurs across model families and zero-shot, reasoning, code, and instruction-following tasks. Similar average accuracy scores do not show how often correct predictions become wrong. Our perturbation–propagation theory factorizes each module's impact into local error and downstream gain along the error's direction. We prove non-asymptotic KL bounds, derive an exact propagation-aware decision radius and a sharp multiclass KL barrier for top-1 changes, and establish guarantees for preserving greedy trajectories; finite-perturbation bounds extend Fisher-based control to every quantile and CVaR. Our theory-guided diagnostic score identifies intervention targets for causal validation. Restoring only the four highest-scoring layers to full precision while keeping all other layers at 3 bits reduces the p99 divergence gap by up to 83%. To improve tail reliability, we introduce two complementary mitigation methods: Propagation-Aware Precision Allocation (PAPA) protects high-impact modules at a fixed bit budget, while Butterfly-Aware Reliability Tuning (BART) learns a lightweight correction at a fixed bit-width. By updating under 2% of parameters, BART raises 3-bit GSM8K accuracy from 0.030 to 0.350. Tail reliability must therefore be measured and reported alongside average performance, and optimized as a quantization objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.