acceptodds
Under review as a conference paper at ICLR 2027

Conformal Prediction-Enabled Sequence-level Policy Optimization for Reasoning Language Models

Abstract

Reinforcement learning improves the reasoning capability of large language models through verifiable rewards for mathematical and coding tasks. In token-level methods such as GRPO, clipping acts on individual token ratios while the reward and advantage apply to a complete response. Recent sequence-level methods align these units but still use fixed clipping bands, whose tail probabilities change as the score distribution evolves. We propose *Conformal Prediction-Enabled Sequence-level Policy Optimization* (CSPO), which adaptively calibrates sequence-level clipping bounds using recent training statistics. The resulting band targets a specified marginal tail budget and adapts to policy drift; calibration conditional on response length requires additional assumptions. We bound the probability that a fresh score falls outside the band by the target miscoverage level plus explicit CDF discrepancy and drift terms. Under additional moment conditions, this bound also controls the objective and gradient error relative to the unclipped GSPO power surrogate. Experiments with Qwen3-1.7B series show improvements over GSPO on difficult mathematical reasoning benchmarks and ALFWorld, with clipping endpoints that adapt during training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.