Group Reward-Mass Optimization for Generative Recommendation
Abstract
Generative recommendation formulates personalized recommendation as a conditional sequence generation task, allowing models to directly generate items aligned with user preferences. Because a user's intent can correspond to multiple valid items, sparse-reward policy optimization can concentrate generation probability on a narrow set of outputs, reducing recommendation diversity and coverage. To address this issue, we propose Group Reward-Mass Optimization (GRMO), a framework that formulates policy learning as direct probability mass matching within candidate groups. GRMO leverages recommendation rewards to construct a target probability mass distribution over sampled candidates. By performing intra-group normalization on both the target and current policy distributions, GRMO minimizes their KL divergence, directly optimizing the relative probability allocation. Experiments on Reddit-v2 show that GRMO improves ranking performance over GRPO while producing more stable, query-conditioned recommendations and preserving access to high-quality alternatives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.