Croissattn: A Complete FP4 Attention Backward
Abstract
FP4 training speeds up the linear layers of language models, but at long context the cost shifts to attention. Four of attention’s six products sit in the backward pass but existing FP4 recipes keep all or part of it in higher precision. CroissAttn extends native FP4 execution to both forward and all four backward attention products. It combines stochastic rounding of gradients and saved-factor recasts with a consistent softmax correction, preserving the quantized forward’s straight- through gradient in conditional expectation. An error decomposition separates noise propagated through the backward pass from new product noise. Query-row- owned accumulation and packed storage reduce the memory required by tiled execution. On long-context workloads, native B200 attention kernels achieve 1.5–2.2× backward speedups over BF16. Training from scratch reaches 10B tokens with validation gaps of 0.0439 nats at 500M parameters and 0.0451 at 1.7B. Native Qwen3-8B continued training reaches a 0.056-nat gap. Under FP4 serving, its checkpoint achieves lower validation loss than directly cast BF16- trained weights. These results connect complete FP4 attention backward to training quality, deployment, and long-context execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.