MOGR: Multi-Objective Generative Recommendation via Objective-Conditioned Distillation and Listwise Generative Policy Optimization
Abstract
Generative recommendation has emerged as a promising paradigm that generates ranked lists in a single autoregressive pass. However, user behaviors such as reads, clicks, and deeper engagement reflect distinct dimensions of user value, and this single-pass design typically forces all objectives into one merged training signal, which exposes inter-objective conflicts: improving one objective often degrades another. To address this, we propose MOGR (Multi-Objective Generative Recommendation), a single-decoder generator that emits, for each list position, a task token declaring which objective the item serves, then generates its Semantic IDs (SIDs). We train MOGR with Objective-Conditioned Distillation (OCD) and Listwise Generative Policy Optimization (LGPO). OCD distills per-objective specialists into one shared decoder under matching objective labels, so each objective keeps its own supervision. LGPO optimizes complete lists under a concave reward, tuning which objective each ranking position serves, a decision OCD alone cannot make. We further introduce OBRQ, an objective-aware codebook variant for industrial logs. On Kwai26, MOGR achieves the best Recall@512 on all three objectives, with 4.3% minimum relative gain (Min-Gain) over MD-CBS and 20.6% over TIGER, and LGPO is compared with a reward-model proxy. Online, OCD+LGPO improves the online Min-Gain over three live objectives relative to frozen OCD by 0.218% under matched exploration (p < 0.05), and by 0.258% under deployed exploration; a separate A/B test gives a 0.524% Min-Gain over the production cascade.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.