CARE-Attn: Component-Aware Residual Error Correction for FP4 Attention QAT
Abstract
Recent advances in attention quantization-aware training (QAT), notably Attn-QAT, have substantially narrowed the quality gap of FP4 attention. Nevertheless, a measurable gap to higher-precision attention remains, and end-to-end recovery alone does not reveal what QAT actually adapts to inside the attention operator. We revisit FP4 attention from a component-wise perspective and find that Attn-QAT adapts highly unevenly: it substantially reduces the error from quantized queries and keys (), while errors from attention probabilities () and values () barely change and a non-negligible residual remains. These residuals also have distinct structures: the residual is sparse and heterogeneous across layers and heads, the error is limited by the quantizer, and the error is associated with channel-wise offsets. We therefore propose CARE-Attn (Component-Aware Residual Error Correction for Attention), which corrects each residual separately: it recomputes predicted high-error tiles in higher precision under adaptive budgets, decouples training and inference precision through -exact QAT, and removes channel offsets from with exactly compensated mean-centering. On video diffusion transformers, CARE-Attn improves generation quality over Attn-QAT, produces outputs closer to those of the BF16 fine-tuned model, and reduces QAT overhead; on language models, it consistently lowers perplexity and substantially improves long-context retrieval. These results show that attention QAT alone is not enough: explicitly correcting the residuals it leaves is key to accurate FP4 attention. Code is available at https://anonymous.4open.science/r/CARE-Attn.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.