QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides
Abstract
Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8. As an emerging low-precision format, NVFP4 combines fine-grained scaling for accuracy preservation with native W4A4 FP4 GEMMs for higher throughput than FP8. However, we find that directly applying NVFP4 to MoE RL rollout is impractical. NVFP4 rollout with BF16 training collapses after roughly 150 steps, accompanied by rapidly growing rollout–trainer log-probability gaps. Through error analysis and operand-isolated experiments, we identify an alignment asymmetry: shared weight QDQ improves cross-engine probability agreement and supports stable RL, whereas activation QDQ lowers the measured agreement and is not helpful to avoid RL collapses. This is because weights can be synchronized across engines, while activations are recomputed online and matching their nominal precision is insufficient for reliable alignment. Therefore, to stabilize NVFP4 RL for MoE, we propose QUantization-error Alignment across Dual Sides (QUADS). On the trainer side, we introduce Asymmetric Quantization-Aware Training fake-quantizing weights while keeping activations unquantized for better alignment. On the rollout side, Residual Activation Compensation corrects high-error activation channels while preserving native W4A4 GEMMs. In our MoE RL experiments on several benchmarks, QUADS achieves BF16-level accuracy, improving average pass@1 by 21.49 points over naive NVFP4 RL, and delivers at least 26% higher rollout throughput than FP8 across the evaluated scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.