acceptodds
Under review as a conference paper at ICLR 2027

Breaking the Even Split: Code Allocation for 4-Bit Post-GELU Quantization in Transformers

Abstract

Static 4-bit post-training quantization leaves each activation tensor only sixteen codes, making their allocation a first-order design choice. In transformer MLPs, GELU confines negative outputs to a narrow, bounded interval, so nearly all of the output range lies on the long, unbounded positive side. Two-branch quantizers fit one range per sign yet split the codes evenly (), spending half of them on the narrow negative interval. Prior methods address this asymmetry but change the code allocation jointly with the grid, region boundaries, or scale search, leaving it without a principled rule or an isolated measurement. We treat code allocation as a design variable and ask whether it can be predicted from unlabeled calibration statistics alone. Minimizing the two branches' uniform-quantization error yields a closed-form rule that sets the allocation from each sign's calibrated mass and range, without labels, gradients, Hessians, or search. Across 84 post-GELU sites, the rule tracks the reconstruction-optimal allocation (Spearman , always within one code) and leave-one-setting-out aggregation recovers a model-wide 4/12 allocation (four negative, twelve nonnegative codes) in all seven settings, the split prior work fixed by design. Held fixed across 19 vision settings (five model families, three resolutions, two datasets), improves top-1 accuracy over the even split in 17, by points on average at identical BitOps, and outperforms promoting three sites to 8 bits at higher cost. Without retuning, it lowers WikiText-2 perplexity versus single-range and even-split quantizers on all six language models tested when only post-GELU activations are reduced to 4 bits. The same allocation at two-sided attention inputs degrades performance in both domains, tying the gain to GELU's bounded negative lobe. Code allocation can thus be read off calibration statistics, and changing it costs no additional BitOps.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.