RelDiff: A Reinforcement Learning Guided Diffusion Framework for Recommendation
Abstract
Diffusion models have demonstrated promising potential in recommender systems. Existing methods typically initiate the generation process by injecting random Gaussian noise, as they have no idea about the authentic user preferences. Owing to the sparse nature of recommendation data, this uninformative paradigm often destroys the original interaction structure and leads to a severe generation collapse. To address the issue, we propose RelDiff, a novel reinforcement learning guided diffusion framework for recommendation. Rather than injecting randomly sampled Gaussian noise, we employ a Preference-Conditional Policy to generate preference noise that actively steers the diffusion process to recover authentic user preferences. Besides, based on a comprehensive analysis of the training dynamics, we design a Negative-aware Refinement strategy to transform the stagnant optimization into an efficient push-pull process. Additionally, we decouple the diffusion process from the inference stage by injecting generative preference into embeddings to maintain training efficiency. Extensive experiments on five real-world datasets demonstrate that RelDiff significantly outperforms state-of-the-art methods. Our work offers an effective solution for enhancing recommendation performance and suggests a novel paradigm for applying reinforcement learning in diffusion frameworks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.