acceptodds
Under review as a conference paper at ICLR 2027

Dual-Curvature Implicit Regularization for GRPO via Sharpness-Aware Minimization

Abstract

Group Relative Policy Optimization (GRPO) has become a widely used algorithm for reinforcement learning with verifiable rewards (RLVR). Common training recipes adopt small learning rates and extended training schedules, while increasing the learning rate can degrade test performance. To investigate which training quantities are associated with generalization performance, we conduct a controlled 20-seed study of GRPO under a fixed training configuration. The strongest associations with test accuracy are observed among the directional Hessian and Fisher curvature metrics, whereas training reward and gradient norm show weak correlations. Motivated by these findings, we propose KL-SAM, a fully first-order optimization algorithm that combines Sharpness-Aware Minimization (SAM)-style perturbations with a local policy-KL regularizer to target objective and policy sensitivity through a single shared perturbation. A local expansion reveals directional Hessian and Fisher contributions, providing an interpretation as implicit dual-curvature regularization. Experiments on models ranging from 1.5B to 7B parameters across multiple objectives and benchmarks demonstrate the effectiveness of our approach. We will release our code upon publication.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.