Focus Calibration for Ultra-Low-Bit All-Reduce in Bandwidth-Constrained Collaborative LLM Inference
Abstract
In bandwidth-constrained environments, collaborative inference of large language models (LLMs) is increasingly bottlenecked by activation communication under tensor parallelism. Although low-bit activation compression can reduce communication overhead in distributed LLM inference, further bit-width reduction is approaching practical limits, as aggressive compression often incurs substantial accuracy degradation or requires costly retraining. To address this challenge, we propose Focus Calibration, a lightweight calibration strategy for retraining-free ultra-low-bit vector quantisation, which enables an index-based all-reduce mechanism that exchanges compact centroid indices instead of full-precision activation tensors. Experimental results show that our method substantially reduces activation communication overhead, achieving up to 38.8% and 12.3% reductions in end-to-end prefill and decoding latency, respectively, while maintaining accuracy comparable to existing quantisation-based approaches at ultra-low bit-widths of 2-2.5 bits. These results establish index-based all-reduce with Focus Calibration as a practical solution for communication-efficient LLM inference in bandwidth-constrained collaborative systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.