acceptodds
Under review as a conference paper at ICLR 2027

FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training

Abstract

Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate the risk induced by the controller that will be deployed; a score calibrated on logged state–action pairs can become miscalibrated after selective action choice; and independent per-resource minimum costs do not in general certify a feasible multi-resource continuation. We introduce FSPO, a feedback-state controller for budgeted LLM RL post-training that addresses these issues jointly. FSPO learns a policy-consistent risk-to-go model whose Bellman target follows the same frozen controller used for future decisions, together with a long-horizon utility model. Decision- conditioned trajectory calibration (DCTC) calibrates risk on cross-fitted trajectories generated by actions selected by provisional controllers. A Pareto resource continuation certificate (PRCC) admits an action only when a non-dominated cumulative reservation remains feasible over the residual horizon. The deployed controller selects the action with the largest pessimistic long-horizon utility among candidates passing both gates. Under a matched GRPO resource envelope, FSPO reaches held-out and OOD accuracy, compared with and for PB2, the strongest evaluated adaptive baseline. Its held-out interval-failure rate is , with training GPU-hours and total development GPU-hours per seed. Three paired training seeds give gains of and percentage points over the contextual bandit on held-out and OOD evaluation. Policy-mismatch, selected-action calibration, and multi-resource continuation tests separately examine the three decision mechanisms.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.