acceptodds
Under review as a conference paper at ICLR 2027

TReQ: Refitting Output Transforms Without Requantizing Language Models

Abstract

Low-bit weight–activation quantization introduces errors that can accumulate across language-model layers. We present TReQ, a post-training quantization method with block-diagonal transformations before and after low-precision matrix multiplication. The output transformations mix channels within small blocks, providing additional reconstruction freedom while keeping quantized weights and scales fixed during each refit. We combine two-sided transformation construction with reconstruction of attention scores, attention context, and downstream projection outputs. The output transformations can be fused into the GEMM epilogue, avoiding a separate kernel and intermediate-memory round trip. Across Ministral-3-8B, Llama-3.1-8B, and Qwen3-8B, TReQ reduces block reconstruction error relative to matched WUSH baselines in INT4, MXFP4, and NVFP4. Under INT4, it lowers task Kullback–Leibler (KL) divergence and improves accuracy on all five benchmarks for every model, while reducing WikiText-2 perplexity. End-to-end improvements in the FP4 formats are less consistent. A fixed-checkpoint Ministral INT4 comparison shows that output refitting alone reduces held-out KL by 8.63%. On NVIDIA B200, fused output transformations add only 0.1–1.5% to MXFP4 full-model prefill latency at the tested workloads with at least 2,048 total tokens.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.