acceptodds
Under review as a conference paper at ICLR 2027

INT8-RL: Training–Inference Co-Optimization for Quantized Rollouts on Ascend

Abstract

Autoregressive rollout generation is a major cost of reinforcement learning (RL) post-training for large language models. INT8 inference can reduce this cost on accelerators with efficient integer kernels, but quantizing the rollout policy changes the distribution from which training trajectories are sampled. Activation outliers further complicate W8A8 execution. We present , a training–inference co-optimization framework for INT8 rollouts on Ascend. The framework combines quantization-aware training with jointly learned smoothing and weight scales, reproducible hash-based rounding, and selective precision fallback. These components adapt the policy to quantized computation while reducing avoidable numerical discrepancies between training and rollout. Experiments on openPangu-7B and Qwen3-30B MoE show substantial reward recovery over the corresponding INT8 baselines. On openPangu-7B with DAPO-Math-17K, the combined W8A8 configuration improves reward from 0.252 to 0.305, compared with 0.309 for BF16. Profiling shows a 19.7% reduction in per-token decode latency. A separate long-sequence ablation that disables training-side fake quantization reduces step time by 18.7%, indicating optimization headroom rather than the measured end-to-end gain of the complete QAT pipeline. These results highlight both the feasibility of INT8 rollouts and the need to control quantization error and training overhead jointly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.