AdaptQuant: Hardware-Aligned Per-Tile MX Format Selection for Post-Training Quantization on NVIDIA Blackwell
Abstract
Post-Training Quantization cuts the cost of serving Large Language Models, and NVIDIA Blackwell's native support for microscaling (MX) formats pushes the speedup further. Yet uniform MXFP4 still incurs substantial quantization error, and existing fixes assign higher-precision formats to sensitive regions using calibration data: a cost repeated for every new model before it can be served at all. We propose AdaptQuant, which assigns an MX element format per tile, the granularity at which Blackwell's block-scaled MMA reads its operands, from a single statistic computed in one pass over the tile. The rule mapping this statistic to a format is derived once, offline, on a single reference model against NVFP4 and reused unchanged on every other model. Across 13 open-weight LLMs on a single RTX 5090, AdaptQuant raises mean perplexity over BF16 by only on WikiText-2 and on C4 while spending fewer effective weight bits on average than the strongest mixed-precision baseline and quantizing faster, cutting stored weight memory by up to over BF16 and reaching BF16's prefill throughput.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.