When Timing Matters: Interleaved Cross-Block Post-Training Quantization
Abstract
Block-wise post-training quantization is making low bit-width compression of large language models increasingly feasible, especially under strict memory budgets. However, PTQ still lacks the global optimization view of quantization-aware training. While most recent PTQ work focuses on block-level improvements, cross-block variants use a moving window to capture inter-block dependencies, offering a step toward global behavior. We go further and propose Interleaved Cross-Block Quantization (ICBQ), a scheduling modification that interleaves short cross-block refinements with block-wise quantization so that errors are corrected before they propagate further, offering a new route toward more global optimization. The interleaved schedule naturally creates boundary overlap, which we analyze separately as a secondary effect. In the reported experiments, ICBQ reduces ternary-quantization perplexity relative to a matched Sequential CBQ baseline, yields finite perplexity in configurations where the baseline has severe degradation, and also transfers to 3-bit and 2-bit GPTQ.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.