acceptodds
Under review as a conference paper at ICLR 2027

Focus Calibration for Ultra-Low-Bit All-Reduce in Bandwidth-Constrained Collaborative LLM Inference

Abstract

In bandwidth-constrained environments, collaborative inference of large language models (LLMs) is increasingly bottlenecked by activation communication under tensor parallelism. Although low-bit activation compression can reduce communication overhead in distributed LLM inference, further bit-width reduction is approaching practical limits, as aggressive compression often incurs substantial accuracy degradation or requires costly retraining. To address this challenge, we propose Focus Calibration, a lightweight calibration strategy for retraining-free ultra-low-bit vector quantisation, which enables an index-based all-reduce mechanism that exchanges compact centroid indices instead of full-precision activation tensors. Experimental results show that our method substantially reduces activation communication overhead, achieving up to 38.8% and 12.3% reductions in end-to-end prefill and decoding latency, respectively, while maintaining accuracy comparable to existing quantisation-based approaches at ultra-low bit-widths of 2-2.5 bits. These results establish index-based all-reduce with Focus Calibration as a practical solution for communication-efficient LLM inference in bandwidth-constrained collaborative systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.