Rethinking Exploration in Generative Recommendation: From More Candidates to Better Candidate Sets
Abstract
Autoregressive generative recommendation uses reinforcement learning (RL) to align sequence generation with recommendation utility. Candidate groups determine experiences for policy learning and choices at inference. However, multi-candidate generation faces a structural exploration–quality trade-off: expanding the search space does not necessarily improve the candidate group. This manifests in two failures. (1) Candidate homogenization: probability-driven decoding repeatedly expands shared high-probability prefixes, spending budget on similar trajectories while leaving alternative decisions underexplored. (2) Candidate-quality degradation: stochastic exploration can discover a few high-reward outcomes while introducing many low-quality candidates, increasing Reward Max but decreasing Reward Mean. Candidate count and maximum reward alone do not characterize effective exploration. These failures motivate learning complementary exploration while controlling quality costs. We propose Q-ROLE (Quality-Gated Role Exploration), a lightweight framework with two coupled components. Role-conditioned generation introduces persistent Anchor and Explorer identities within a shared policy, making task-focused exploitation and complementary exploration learnable responsibilities. A quality-gated coverage reward grants a bonus for divergence from the Anchor only when an Explorer meets an Anchor-relative base-reward threshold; otherwise, it receives only its base reward. The learned roles are retained during inference, carrying this division of responsibilities into candidate generation. Q-ROLE preserves the model backbone and requires no auxiliary prefix-value model. Extensive experiments on public recommendation benchmarks demonstrate that Q-ROLE improves a range of strong recommendation baselines, delivering higher recommendation accuracy under matched candidate budgets, while online experiments in an industrial generative reranking system report improved ranking performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.