Generative Visual Recommendation through Retrieval-to-Generation Mapping
Abstract
Generative recommendation has evolved along two major paradigms: generative retrieval (GenRet) and personalized visual generation (PVGen). Specifically, GenRet captures collaborative signals from user–item interactions but remains restricted to existing catalog items. While PVGen synthesize new items, yet may not fully exploit these signals for user preference modeling. However, integrating the two paradigms presents challenges in computational efficiency and semantic misaligned between their retrieval and visual generation latent spaces. To address these challenges, we propose Retrieval-to-Generation Mapping (R2G-Map), which efficiently and effectively integrates GenRet and PVGen by aligning recommendation-oriented preference representations with the visual generation space. Specifically, R2G-Map uses continuous recommender hidden states before discrete token selection to retain fine-grained collaborative and predictive information. To bridge the two latent spaces, we introduce a Factorized Feature Predictor that predicts content, description, and collaborative features and constructs visual condition for image generation. We further develop a Dual-Branch Visual Generator that combines this condition with preference representations through joint attention between an image branch and a latent branch. The resulting visual-token distributions support catalog ranking, while the corresponding token sequences are decoded into personalized item images. Experiments on bundle and sequential recommendation demonstrate that R2G-Map improves ranking performance while generating visually coherent and compatible item images.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.