acceptodds
Under review as a conference paper at ICLR 2027

ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization

Abstract

The alignment of Large Language Models (LLMs) increasingly relies on Reinforcement Learning from AI Feedback (RLAIF) for non-verifiable domains such as long-form question answering and open-ended instruction following. These domains often use LLM-based auto-raters to provide granular, multi-tiered discrete rewards (e.g., – rubrics) that can be stochastic due to prompt sensitivity and sampling randomness. We consider the black-box setting in which the optimizer observes only the returned discrete reward and does not assume access to the auto-rater's parameters, logits, hidden states, or auxiliary confidence estimates. We empirically characterize this stochasticity across multiple auto-rater configurations. Such variability enters standard advantage estimators such as GRPO and MaxRL through group-normalization statistics, allowing a noisy reward sample to alter the update assigned to the entire group. Repeatedly sampling and aggregating rewards may reduce this noise, but doing so substantially increases evaluation cost. To address this bottleneck, we introduce rdinal ecomposition for obust olicy ptimization (), a framework that decomposes discrete rewards into a sequence of ordinal binary indicators. By independently computing and accumulating advantages across progressively challenging success thresholds, ODRPO localizes score variability to the thresholds crossed by a score change while supporting flexible weighting schemes and an implicit ordinal curriculum. Empirically, across GRPO, MaxRL, and OTB, three policy models, and seven evaluation benchmarks, all main-experiment configurations ODRPO and its variance-aware variant achieve positive mean relative gains ranging from to . Ablations over auto-rater temperature and reward granularity further show consistent improvements over GRPO, with finer reward resolution producing the clearest trend. These gains incur no material per-step training overhead. ODRPO further attains the noise robustness sought by repeated auto-rater sampling, surpassing MaxRL with -sample majority voting on AlpacaEval while using a single reward sample. Finally, our theoretical analysis establishes that the ODRPO formulation retains a well-defined global scalar objective in arbitrary discrete reward spaces when built on estimators with valid binary objectives.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.