acceptodds
Under review as a conference paper at ICLR 2027

Ratio-Free Policy Optimization: A Bias-Variance View of On-Policy Fine-Tuning

Abstract

Practical LLM reinforcement learning is never strictly on-policy: each batch of rollouts is reused for many optimizer steps, and rollouts come from a separate inference engine. We show that GRPO degrades under both drifts — and on coding collapses — and explain why. We formalize on-policy fine-tuning (OPFT), updating a model on its own rollouts, as the OPFT loop, of which GRPO and iterative rejection-sampling SFT are realizations, and propose RFPO (Ratio-Free Policy Optimization), a novel realization that optimizes the surrogate — a reward-weighted maximum-likelihood objective requiring no importance ratio — as a local proxy for the on-policy expected reward . We identify three mechanisms — z-score advantages, negative samples, and trust-region clipping — that make this surrogate stable, each with a distinct role verified by theory and ablations; the analysis also explains the well-known collapse of iterative rejection-sampling SFT. GRPO realises analogous mechanisms through different machinery: the core updates of RFPO and GRPO coincide at strict on-policy and diverge under the practical sampling gaps that any modern LLM RL stack faces (multiple optimizer steps per batch; decoupled rollout/training engines). A bias-variance analysis predicts that GRPO's importance-sampling variance scales with the gap while RFPO's bias is controlled by a per-token trust region, so RFPO dominates in MSE under sufficient drift. On Qwen3-4B training on POLARIS/DeepCoder we verify the prediction along both drift axes: as we increase optimizer steps per batch, GRPO degrades from 0.90 (MATH-500)/0.67 (LCBv6) at 1 step through 0.87/0.38 at 8 steps to 0.85/0.24 at 32 steps, while RFPO stays at 0.90/0.67; under bypass-mode rollout drift (8 steps), GRPO collapses () while RFPO is nearly unaffected (0.90/0.67) — and, needing no reference model, KL term or importance ratio, RFPO trains 1.3–1.9 faster per step.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.