acceptodds
Under review as a conference paper at ICLR 2027

AdaPivot: An Adaptive Core-Tail Representation for Low-Bit LLM Quantization

Abstract

Post-training quantization (PTQ) of Large Language Models (LLMs) often suffers substantial accuracy degradation because fixed numerical formats poorly accommodate the pronounced outliers in weights and activations. Through extensive profiling across diverse models, we observe a distributional structure: most values concentrate in a compact core, while the outliers, which account for only a small fraction of the values, form a sparse tail, typically separated from the core by a sharp magnitude transition and a nearly unoccupied interval. Building on this observation, we introduce AdaPivot, an adaptive low-bit numerical format that partitions the representable range into distinct core and tail regions. AdaPivot independently adapts the representation granularity of the two regions, together with their boundary and dynamic ranges, to match the underlying tensor distribution within a compact parametric form. We further formulate format fitting as a joint discrete–continuous reconstruction problem and develop a tailored solver that efficiently co-optimizes discrete level selection and continuous format parameters. An integer-constrained variant further reduces metadata overhead while preserving the adaptive structure. Experiments on various models from 7B to 70B parameters show that AdaPivot consistently outperforms representative PTQ baselines. Under W4A4 quantization, it improves average zero-shot accuracy by up to 7.5% over the state-of-the-art baseline, while W4A8 performance closely approaches the FP16 baselines. Our code is available at https://anonymous.4open.science/r/AdaPivot-A.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.