acceptodds
Under review as a conference paper at ICLR 2027

Full-Stack NVFP4: Stable, Role-Aware 4-Bit LLM Pretraining Beyond Linear Projections

Abstract

NVFP4 pretraining typically targets Transformer projections while leaving optimizer states, optimizer computation, and attention in higher precision. Extending NVFP4 beyond projections is not a uniform quantization problem: each module has a distinct numerical role and failure mode. We present , a role-aware framework instantiated by four independently composable recipes. retains a compact BF16 projection subspace within full-shape NVFP4 computation; gradient decoupling and periodic SVD realignment stabilize its low-rank factors, reducing the linear-only loss gap from to . transforms persistent momentum states before storage, stabilizes direct NVFP4 Newton–Schulz iterations, and protects softmax-sensitive products in BF16. On 3B pretraining with 64B tokens, BF16 and Full-Stack NVFP4 reach losses of and , a gap, with similar aggregate zero-shot results. Native four-block measurements on one RTX 5090 show 2.50–2.83 Root speedups over optimized BF16 and 37.9–42.5% lower AdamW peak memory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.