acceptodds
Under review as a conference paper at ICLR 2027

Pinpoint Quantization: A Mechanistic Account of Tail-Capability Failures in Quantized LLMs

Abstract

Weight quantization is essential for deploying large language models (LLMs), yet quantization below bits can sharply degrade accuracy. Prior work has primarily characterized this degradation through model behavior or weight statistics, offering limited insight into its internal causes or how to enable targeted mitigation. Instead, we provide a *mechanistic account* of this failure by tracing capability-specific degradation to internal model components. We find that the damage is strikingly selective: it disproportionately impairs ”tail” capabilities (e.g., niche languages, long-context processing, complex factual queries, and multi-hop reasoning) while largely preserving ”head” capabilities. To explain this selectivity, we analyze module function and discover two predominant perturbation mechanisms: attention heads undergo *semantic direction shifts*, whereas MLP neurons undergo *activation-magnitude suppression*. Crucially, the mechanistic fragility is concentrated in a sparse set of key modules: restoring their weight precision recovers most of the accuracy loss on the degraded tasks, suggesting that precise bit allocation, rather than bit width alone, is critical to sub--bit quantization. As a practical application of these findings, we introduce Pinpoint Quantization, which uses causal-effect scores to guide group-knapsack bit allocation. Across benchmarks and dense and MoE models ranging from B to B parameters, it achieves near-lossless accuracy at an average weight precision below bits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.