acceptodds
Under review as a conference paper at ICLR 2027

CoBIT: Unified Optimization of Codebooks, Assignments, and Activation Composition for Extreme Low-Bit LLM Quantization

Abstract

Post-training quantization is essential for efficient large model inference, yet preserving accuracy at ultra-low budget below 4 bits remains challenging. Clustering-based quantization provides adaptive non-uniform representations, but existing methods use limited information to construct codebooks and assign weights, while activation quantization is typically handled separately. We present CoBIT, a post-training quantization framework that optimizes weight clustering and activation quantization under a common layer-output reconstruction objective. For weights, CoBIT alternates Hessian-aware codebook construction with error-compensated assignment. For activations, CoBIT selects a per-token exponent–mantissa composition based on layer-output error. We introduce PhaseLUT to replace runtime composition search with a lookup, and a hardware-aware kernel that integrates composition selection, activation quantization, and low-precision computation. Across all four general-purpose models, CoBIT improves average accuracy over the strongest baseline in each setting by 4.8% and 35.9% points with 4-bit activations and 3- and 2-bit weight clustering, respectively. With 4-bit weight indices, it also achieves up to end-to-end decoding speedup over dense FP16.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.