Does Quantization Preserve How Language Models Answer? Causal-Equivariant Quantization for Symbol-Binding Circuits
Abstract
Post-training quantization (PTQ) is typically optimized to preserve weights and activations while maintaining perplexity, or task accuracy. However, these objectives do not reveal whether quantization preserves the causal computation underlying model decisions. Recent studies have localized answer selection and symbol binding in formatted multiple-choice question answering (MCQA) to a sparse set of attention components, making it possible to directly examine their preservation under quantization. To preserve the causal mechanisms underlying formatted MCQA after quantization, we propose Causal-Equivariant Quantization (CEQ), a structured mixed-precision framework consisting of three components: counterfactual causal probing, causal-regret-based precision allocation, and causal-equivariant calibration. First, counterfactual causal probing identifies model components whose causal behavior is most affected by low-bit quantization across different answer formats. CEQ measures these changes using causal-effect distortion (CED) and format-equivariance error (FEE). Second, causal-regret-based precision allocation uses this information to assign higher precision to causally important and quantization-sensitive components under a global bit budget. Third, causal-equivariant calibration (CEQ-Cal) further reduces the discrepancy between the quantized and full-precision models while maintaining consistent behavior across answer formats by preserving output distributions, answer margins, causal effects, and format equivariance. We evaluated CEQ on a novel controlled symbol-binding dataset constructed specifically for this work and on five public MCQA datasets. On the controlled symbol-binding benchmark, CEQ achieves 83.40 ± 14.21% worst-format accuracy at 2.72 effective matrix bits versus 68.00 ± 20.64% for uniform 2-bit RTN. Across the five public MCQA datasets, CEQ consistently improves the worst-format accuracy over the matched TaCQ baseline, achieving an average gain of 9.86% points. The effectiveness of CEQ further generalizes across model scales and architecture families, achieving 76.55% MMLU accuracy on Qwen2.5-32B-Instruct at 2.12 bits and a 67.95% zero-shot average on Llama-3-8B at 4.08 bits. These results show that preserving prediction quality under quantization does not necessarily preserve the underlying causal computation, and motivate mechanism preservation as a distinct objective for low-bit language-model compression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.