acceptodds
Under review as a conference paper at ICLR 2027

When Mismatch Misleads: Stabilizing Low-Precision Reinforcement Learning via Reliable Correction

Abstract

Reinforcement learning (RL) has become a key component of large language model post-training, but its substantial computational cost increasingly motivates end-to-end low-precision training. Despite improving training efficiency, it also introduces numerical perturbations that amplify train-inference mismatch and create a tradeoff between efficiency and stability. Existing approaches primarily stabilize low-precision RL by aligning rollout and actor policies. However, actor-rollout alignment alone does not guarantee improvement of the underlying full-precision objective, since the observed mismatch conflates genuine policy mismatch with numerical noise. We therefore formulate low-precision RL as a three-policy alignment problem, where noisy rollout and actor policies provide observable training signals while the inaccessible full-precision policy defines the learning problem. Based on this formulation, we propose **ERIS (Excess-Risk-guided Importance Sampling)**, an excess-risk-guided method that corrects high-reliability mismatch tokens while retaining identity weighting for noise-dominated tokens. Across multiple Qwen model scales and architectures, ERIS consistently improves FP8 training stability and performance. ERIS improves the average mathematical reasoning score over standard FP8 by 14.29% on Qwen3-8B, while achieving performance comparable to BF16. Mechanism analysis further confirms the advantage of three-policy alignment in preserving consistency with the full-precision policy. Meanwhile, ERIS preserves the efficiency benefits of FP8, achieving up to a 1.17 end-to-end throughput improvement over BF16.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.