acceptodds
Under review as a conference paper at ICLR 2027

CoNestQ: Coordinated Nested Quantization for Low-Bit LLM Deployment

Abstract

Low-latency and data-locality requirements are shifting LLM inference from centralized clouds toward distributed, multi-tier architectures such as the Intelligence Delivery Network (IDN), which distributes intelligence across cloud, regional, edge, and local nodes on demand. Serving such heterogeneous nodes economically requires a single model representation from which multiple precision levels can be derived. Fixed-precision quantization cannot meet diverse node budgets without storing a separate model per bit width, while existing nested methods leave the cross-bit coupling induced by representation sharing largely unaddressed, with several bit-widths unusable on some models. To address these limitations, we propose CoNestQ (Coordinated Nested Quantization), a post-training quantization framework that derives all target precisions from a shared parent representation and coordinates the three coupled stages that prior nested methods address only partially: (i)joint quantization-parameter calibration across target bit-widths, (ii)balanced parent-code selection guided by normalized excess error, and (iii)branch-wise error compensation with periodic cross-precision synchronization. Across the full 3–8-bit range on four LLMs, CoNestQ attains a 100% validity rate with strong low-bit robustness. Relative to the non-linear nested baseline, it reduces multi-precision storage by 26–34%, lowers peak GPU memory by about 38% on average, and improves inference throughput by about 16%; relative to six independently quantized models, multi-precision storage shrinks by 5.5–6.2.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.