acceptodds
Under review as a conference paper at ICLR 2027

AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation

Abstract

Hadamard transforms have become a standard tool for stabilizing MXFP4 LLM training by spreading outliers within its 32-element scaling blocks. We show that their effectiveness depends on a simple geometric factor: how outliers are oriented relative to the scaling axis. Outliers perpendicular to the axis create within-block concentration that Hadamard mixing can effectively reduce, whereas outliers aligned with the axis enlarge the shared block scales, leaving a scale-driven component that rotation does not directly address. We call this the scaling-axis blind spot. Measurements of LLM training tensors reveal stable Row-wise, Column-wise, and None outlier patterns whose combinations vary across GEMMs; notably, 97.8–100% of weight-gradient GEMMs contain a blind-spot operand. We introduce AdaHOP, an outlier-pattern-aware MXFP4 training framework that identifies these patterns in a short calibration phase, uses Inner Hadamard Transform (IHT) as the low-precision backbone, and selectively extracts only the blind-spot outlier rows or columns into a small BF16 path. With fused Triton kernels, AdaHOP-Lv1 executes less than 1% of GEMM FLOPs in BF16 while trailing BF16 by only 0.23 points in average zero-shot accuracy across models from 1B to 8B. On Llama3.1-8B, it reduces memory by and accelerates end-to-end training by over BF16.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.