Conformal Prediction-Enabled Sequence-level Policy Optimization for Reasoning Language Models
Abstract
Reinforcement learning improves the reasoning capability of large language models through verifiable rewards for mathematical and coding tasks. In token-level methods such as GRPO, clipping acts on individual token ratios while the reward and advantage apply to a complete response. Recent sequence-level methods align these units but still use fixed clipping bands, whose tail probabilities change as the score distribution evolves. We propose *Conformal Prediction-Enabled Sequence-level Policy Optimization* (CSPO), which adaptively calibrates sequence-level clipping bounds using recent training statistics. The resulting band targets a specified marginal tail budget and adapts to policy drift; calibration conditional on response length requires additional assumptions. We bound the probability that a fresh score falls outside the band by the target miscoverage level plus explicit CDF discrepancy and drift terms. Under additional moment conditions, this bound also controls the objective and gradient error relative to the unclipped GSPO power surrogate. Experiments with Qwen3-1.7B series show improvements over GSPO on difficult mathematical reasoning benchmarks and ALFWorld, with clipping endpoints that adapt during training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.