acceptodds
Under review as a conference paper at ICLR 2027

Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

Abstract

Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, *i.e.*, the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose **Group-Marginalized Advantage Estimation (GMAE)**, which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.