acceptodds
Under review as a conference paper at ICLR 2027

Rollout-Guided QAT for FP4 Reinforcement Learning of MoE Language Models

Abstract

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantiza- tion accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose NTR, an FP4 quantization framework for RL training of Mixture-of- Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. NTR incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, NTR adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate NTR on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that NTR enables joint FP4 weight/activation and FP4 KV- cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4× rollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.