acceptodds
Under review as a conference paper at ICLR 2027

ARGUS: Seeing Beyond the Kernel in Stage-Aware NVFP4 Attention

Abstract

NVFP4 attention leverages high-throughput FP4 arithmetic on NVIDIA Blackwell GPUs, but existing training-free methods largely optimize the main fused kernel and apply similar low-precision execution across inference stages. This kernel-centric focus overlooks two important bottlenecks: preprocessing and cache construction can constitute a substantial fraction of prefill operator latency, whereas autoregressive decoding is primarily constrained by streaming the growing KV cache. We introduce **ARGUS**, a stage-aware NVFP4 attention design that looks beyond kernel-level acceleration and places low precision where each inference stage benefits from it most. This stage-aware design is realized through two complementary components. **Cache-Ready Prefill** replaces input-dependent statistics with **Offline Sparse Centering**, reducing runtime preprocessing while directly materializing a persistent NVFP4 KV cache. **Cache-Native Decoding** directly consumes this compressed cache with BF16 attention, decoupling cache precision from compute precision to reduce memory traffic without unnecessary low-precision transformations. Across reasoning and long-context evaluations on GQA models, ARGUS maintains competitive model quality and achieves up to **4.09×** prefill and **2.26×** decode operator-level speedups, yielding up to **2.40×** end-to-end acceleration for the evaluated 128K-input workload on SGLang.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.