ARGUS: Seeing Beyond the Kernel in Stage-Aware NVFP4 Attention
Abstract
NVFP4 attention leverages high-throughput FP4 arithmetic on NVIDIA Blackwell GPUs, but existing training-free methods largely optimize the main fused kernel and apply similar low-precision execution across inference stages. This kernel-centric focus overlooks two important bottlenecks: preprocessing and cache construction can constitute a substantial fraction of prefill operator latency, whereas autoregressive decoding is primarily constrained by streaming the growing KV cache. We introduce **ARGUS**, a stage-aware NVFP4 attention design that looks beyond kernel-level acceleration and places low precision where each inference stage benefits from it most. This stage-aware design is realized through two complementary components. **Cache-Ready Prefill** replaces input-dependent statistics with **Offline Sparse Centering**, reducing runtime preprocessing while directly materializing a persistent NVFP4 KV cache. **Cache-Native Decoding** directly consumes this compressed cache with BF16 attention, decoupling cache precision from compute precision to reduce memory traffic without unnecessary low-precision transformations. Across reasoning and long-context evaluations on GQA models, ARGUS maintains competitive model quality and achieves up to **4.09×** prefill and **2.26×** decode operator-level speedups, yielding up to **2.40×** end-to-end acceleration for the evaluated 128K-input workload on SGLang.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.