acceptodds
Under review as a conference paper at ICLR 2027

HyperHadamard: Layout-Aware Hadamard Transforms on Tensor Cores

Abstract

Low-bit quantization is an appealing way to cut the memory and bandwidth cost of large language models. But activation outliers make few-bit quantization unstable, and rotation-based methods address them by mixing channels before quantization. The Hadamard transform is the rotation of choice because it admits a fast algorithm. Tensor-core FWHT kernels made these rotations cheap enough to deploy, but at large rotation dimensions existing kernels sustain only 42–77% of achievable memory bandwidth, even though the transform is memory-bound. We identify intermediate-layout materialization as a source of excess instructions and dependency stalls. We propose HyperHadamard, which treats butterfly stages as rotations along the binary axes of a hypercube and assigns those axes to tensor-core register layouts at compile time, so intermediate transposes become metadata rather than data movement. At rotation dimensions large enough for these overheads to dominate, HyperHadamard is up to 1.59 faster than HadaCore, the tensor-core FWHT used in vLLM, and stays within 1.7% of achievable memory bandwidth. Together with a gather remap and a fused sign load, a true-FP32 instantiation of the same schedule speeds up a billion-dimensional randomized gradient projection used for training-data selection by 1.75 while remaining bit-exact. The same design supports non-power-of-two transforms built from available catalog Hadamard matrices, whereas prior fast kernels expose only power-of-two and selected fixed-factor sizes. Integrated into vLLM, a Llama-3.1-8B QuIP W4A16 model produced by llm-compressor runs 14.7 faster end to end than vLLM's dense GEMM path and 1.9 faster than the same factorization assembled from Dao-AILab kernels and cuBLAS, with no measurable change in perplexity or zero-shot accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.