TReQ: Refitting Output Transforms Without Requantizing Language Models
Abstract
Low-bit weight–activation quantization introduces errors that can accumulate across language-model layers. We present TReQ, a post-training quantization method with block-diagonal transformations before and after low-precision matrix multiplication. The output transformations mix channels within small blocks, providing additional reconstruction freedom while keeping quantized weights and scales fixed during each refit. We combine two-sided transformation construction with reconstruction of attention scores, attention context, and downstream projection outputs. The output transformations can be fused into the GEMM epilogue, avoiding a separate kernel and intermediate-memory round trip. Across Ministral-3-8B, Llama-3.1-8B, and Qwen3-8B, TReQ reduces block reconstruction error relative to matched WUSH baselines in INT4, MXFP4, and NVFP4. Under INT4, it lowers task Kullback–Leibler (KL) divergence and improves accuracy on all five benchmarks for every model, while reducing WikiText-2 perplexity. End-to-end improvements in the FP4 formats are less consistent. A fixed-checkpoint Ministral INT4 comparison shows that output refitting alone reduces held-out KL by 8.63%. On NVIDIA B200, fused output transformations add only 0.1–1.5% to MXFP4 full-model prefill latency at the tested workloads with at least 2,048 total tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.