acceptodds
Under review as a conference paper at ICLR 2027

Where Rounding Happens Matters: Rounding-Location Alignment Stabilizes LLM Reinforcement Learning

Abstract

Reinforcement learning for large language models often uses separate systems for inference and training. Even with identical weights and BF16 tensors, implementation differences can make an intended on-policy update off-policy and destabilize training. Existing remedies align the entire inference and training stacks, change global precision, or modify the learning algorithm. We ask whether aligning selected parts of the execution stack can stabilize BF16 RL while other implementation differences remain. We align training-side rounding locations with inference in three regions—RoPE, SwiGLU, and residual RMSNorm—by retaining selected intermediates in FP32 until their BF16 outputs. Across four Qwen-family models, aligning all three regions prevents the collapse observed under the default BF16 baseline; RoPE-path alignment alone also suffices in the studied runs, despite remaining numerical discrepancies. The interventions keep the inference engine, matrix-multiplication and attention kernels, and learning objective unchanged.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.