Quantar: Quantizing All-Reduce for Efficient Tensor-Parallel Serving
Abstract
Quantization has reached nearly every tensor a language model stores or computes, yet tensor-parallel all-reduce still communicates in BF16. We find that this communication accounts for as much as 37% of prefill device time across models ranging from an 8B dense transformer to a 671B mixture-of-experts model, and its share grows as lower-precision computation makes the rest of inference faster. We introduce Quantar to quantize the reduction itself. Existing dynamic quantization, if applied to all-reduce, would require an extra communication round for scale agreement; without it, the reduced values cannot be mapped back to the correct sum. Quantar instead freezes communication scales calibrated once and shares them across ranks and requests. We show that transformer normalization places an input-independent bound on all-reduce activations, making one frozen scale practical across ranks. On eight B300 GPUs, Quantar speeds up the complete communication boundary by up to 1.81× and prefill by up to 1.28×, and on eight RTX PRO 6000 GPUs, performance measurements reach 1.92× and 1.61×. At matched settings, Quantar outperforms per-hop FP8 by up to 1.24× at the communication boundary and 1.10× in prefill, while avoiding its 1.54× error growth from two to eight ranks. We also show that quality remains at BF16 parity: full 12,032-question MMLU-Pro scores change by at most 0.42 percentage points across five models, and all nine benchmarks and 33 subtasks evaluated across three models remain within 2σ. Quantar also matches BF16 on LongBench-E, GSM8K-CoT, and RULER from 8k through 128k, including 85.4 versus 85.0 at 128k. Quantar's advantage grows as models become larger, more sharded, and more heavily quantized, extending low-bit gains across the inference path.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.