Role-Aware Quantization for Reliable Transformer Inference
Abstract
Quantization is widely used to reduce Transformer memory usage and accelerate inference. Prior work has commonly used uniform activation quantization, assigning the same bit-width across activations. However, Transformer components perform distinct functions, including query–key routing, value transport, feed-forward computation, and normalization. We show that these functional roles differ in sensitivity to quantization error and therefore require different precision levels to preserve predictions. To allocate precision efficiently under a fixed budget, we introduce Role-Aware Reliability (RAR) allocation. RAR quantizes one layer–role activation site at a time and measures the resulting change in the next-token distribution using Kullback–Leibler (KL) divergence. The resulting functional risk atlas pinpoints where low precision most strongly distorts predictions. Building the functional risk atlas for a 7B checkpoint requires only a small sample of calibration text and approximately one to two minutes. Our allocation rule, RAR-log, uses the atlas to assign higher precision to high-risk sites and lower precision to low-risk sites under a fixed cost-weighted average bit-width budget. By accounting for how quantization errors in different roles affect predictions, RAR complements SmoothQuant-style activation smoothing. Our theoretical analysis explains how quantization errors from different sites combine and proves bounds on the resulting changes in model predictions. With the same target precision budget, the combined policy, Smooth-RAR-log, removes an average of 94.1% of the language-modeling degradation caused by uniform 6-bit activations, compared with 87.4% for smoothing alone. The combined policy improves on smoothing alone in all 24 paired model–corpus evaluations, covering three corpora and eight Transformer checkpoints ranging from 124M to 7B parameters. Even with packed activations and 8-bit quantized weights, Smooth-RAR-log lowers next-token prediction loss compared with smoothing alone across the evaluated models. Smooth-RAR-log also improves average accuracy on multiple-choice benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.