acceptodds
Under review as a conference paper at ICLR 2027

AWSLQ: Activation-Wise Sparse-Ladder Quantization for Accurate W4A4 Single-Load Decoding

Abstract

On FP4-native GPUs, deployed serving stacks already run W4A4 — FP4 weights times FP4 activations — at decode, yet the quantization literature retreats to 16-bit decode because W4A4 is where accuracy collapses. Small-batch decode is bound by streaming the weights, so its tensor cores sit mostly idle. AWSLQ, a training-free W4A4 scheme that requires no external calibration corpus, spends that idle compute: it multiplies each weight tile by several FP4 activation terms instead of one, here a ladder of N structured-sparse NVFP4 operands, peeled online, that share one weight load. Two-operand ablations had concluded that more sparse computation does not recover the loss; the knee is at N=3, and by N≈4 the ladder is on the W4A16 floor, to within +0.005 perplexity on ten of eleven models. A selective reverse-smoothing fold, in the direction SmoothQuant advises against, lowers the weight-only floor itself; the ladder absorbs the activation outliers the fold creates and lands on the lowered floor, while a single dense term stays above it; whether to fold is decided per model from its own generated text at preparation time. A fused Blackwell kernel runs the ladder as one M-stacked sparse GEMM into a single accumulator and beats cuBLASLt's FP4 GEMM across a Llama-3.1-8B decode block at M=16–64 (1.26× to 1.07× on an RTX 5090, 1.08× at M=16 on a B200). AWSLQ matches or beats every published W4A4 method, calibrated or not, on 16 of the 18 model–dataset cells with a published comparator, at coarser block-32 scales, and on GSM8K it matches the accuracy of retreating to 16-bit decode activations. The accuracy comes from the number of FP4 terms that share one weight load. A sparse rung performs half the multiply-accumulates of a dense term, so the ladder also offers an intermediate arithmetic budget, three rungs at one and a half dense terms, which chooses the refined half of each token's channels online.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.