Don't Let Conflicting Rewards Pull Generative Recommendations Off Target
Abstract
LLM-based generative recommender systems must satisfy multiple objectives, but in practice these objectives may have different priorities, with recommendation accuracy serving as the primary objective and auxiliary objectives such as diversity and rationale coherence providing complementary training signals. In widely adopted reinforcement learning with verifiable rewards (RLVR) frameworks such as Group Relative Policy Optimization (GRPO), these heterogeneous rewards are merged into a fixed weighted sum, implicitly treating all rewards as equally tradable. When rewards conflict and their signals differ in scale, difficulty, and saturation, fixed aggregation can cancel gradients or let fast-improving auxiliary rewards dominate, resulting in reward hacking and signal collapse that can even degrade recommendation accuracy. To address this, building on the RLVR-style optimization framework, we propose Groupwise Gradient-Aligned Reward-Decoupled Policy Optimization (RPO), which dynamically selects reward-aggregation weights to maximize alignment with the primary accuracy gradient while constraining the resulting update to avoid substantial first-order degradation of auxiliary objectives. This yields a primary-centered update within the region defined by auxiliary compatibility constraints. Concretely, RPO computes per-reward aggregation weights via a simplex linear program, which identifies a weighted combination of the primary (i.e., accuracy) and auxiliary reward (i.e., diversity) gradients that is maximally aligned with the primary reward gradient, while enforcing that auxiliary gradients are non-adversarial to the resulting update direction. Experiments on five Amazon domains and MovieLens show generally consistent gains in accuracy, diversity, and rationale coherence. Beyond offline evaluation, our method is deployed in production and achieves the highest click-through rate in a real-world online A/B test.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.