INT8-RL: Training–Inference Co-Optimization for Quantized Rollouts on Ascend
Abstract
Autoregressive rollout generation is a major cost of reinforcement learning (RL) post-training for large language models. INT8 inference can reduce this cost on accelerators with efficient integer kernels, but quantizing the rollout policy changes the distribution from which training trajectories are sampled. Activation outliers further complicate W8A8 execution. We present , a training–inference co-optimization framework for INT8 rollouts on Ascend. The framework combines quantization-aware training with jointly learned smoothing and weight scales, reproducible hash-based rounding, and selective precision fallback. These components adapt the policy to quantized computation while reducing avoidable numerical discrepancies between training and rollout. Experiments on openPangu-7B and Qwen3-30B MoE show substantial reward recovery over the corresponding INT8 baselines. On openPangu-7B with DAPO-Math-17K, the combined W8A8 configuration improves reward from 0.252 to 0.305, compared with 0.309 for BF16. Profiling shows a 19.7% reduction in per-token decode latency. A separate long-sequence ablation that disables training-side fake quantization reduces step time by 18.7%, indicating optimization headroom rather than the measured end-to-end gain of the complete QAT pipeline. These results highlight both the feasibility of INT8 rollouts and the need to control quantization error and training overhead jointly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.