When Negative Advantages Are Unreliable: Risk-Calibrated Negative Suppression for Generative Recommendation
Abstract
In generative recommendation, reinforcement learning (RL) optimizes recommendation policies using rewards derived from user feedback. Methods such as GRPO compute advantages from the relative rewards of multiple recommendations generated for the same context and suppress those with negative advantages. However, the user feedback available during training is incomplete: recommendations unsupported by observed feedback may still match user interests, so suppressing them risks discouraging otherwise plausible recommendations. To investigate this problem, we measure the fraction of negative-advantage recommendations whose advantages become positive when evaluated against subsequent feedback. Across three recommendation domains, 5.2%–13.1% of recommendations with negative observed advantages receive positive advantages under validation feedback, demonstrating the recurrence of these sign conflicts. This motivates selectively attenuating feedback-sensitive negative updates while retaining useful negative learning signals. The challenge is to determine which updates to attenuate using information available during training. We find that sign conflicts are more likely among negative-advantage recommendations with better observed retrieval ranks, providing an observable cue available during training. Building on this insight, we propose Risk-Calibrated Negative Suppression (RCNS), which reuses this ordering to allocate continuous attenuation without reversing advantage signs. Using held-out validation feedback, we calibrate a shared domain-level attenuation parameter to maximize coverage of sign-conflicting advantage mass subject to an empirical risk budget on attenuating non-conflicting negative mass. The calibrated parameter remains fixed during policy training and requires no subsequent feedback for newly generated recommendations. Offline experiments show that RCNS achieves greater held-out advantage-weighted coverage of sign-conflicting updates than uniform attenuation under the same validation risk budget. In the main downstream comparison, RCNS improves mean HR@5 across five independent seeds over GRPO by 5.0%–19.7% across three domains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.