acceptodds
Under review as a conference paper at ICLR 2027

Analysis of MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible error

Abstract

MXFP4 arithmetic can dramatically accelerate reinforcement learning (RL) post-training of large language models (LLMs), yet the quantization error introduces severe accuracy degradation. Existing work treats the quantization error as a monolithic noise term, missing the distinct mechanisms upon interpreting how quantization error damages training. We prove that this error is not monolithic but admits an exact, privileged decomposition into three additive components: scale bias from power-of-two rounding, deadzone truncation from zeroing small values, and grid noise from rounding to the 4-bit grid. The deadzone component is exactly orthogonal to the other two, and, crucially, the grid component is invariant to scale precision: refining the block scale reduces scale bias without touching grid noise, driving the total error toward an irreducible floor set by the grid and deadzone alone. Each component dominates a distinct RL failure mode: scale bias accumulates multiplicatively through the backward pass, affecting gradient accuracy; deadzone truncation degrades rollout quality; and grid noise raises the policy's entropy. Our contribution is this diagnostic framework rather than a new correction: to validate it, we use three existing, mechanism-targeted techniques as controlled interventions (Macro-block scaling for scale bias, Outlier Fallback for deadzone, and Adaptive Quantization Noise for grid noise/entropy), each removing one component and recovering its predicted RL pathway. On Qwen2.5-3B (dense) this recovers BF16 accuracy to within 0.7 pp, on Qwen3-30B-A3B-Base (MoE) to parity, and on Llama-3.1-8B (a different family) to within 3.2 pp. The same component-to-pathway mapping holds across all three, and the results generalize from GSM8K to MATH500.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.