Mitigating Sampling Imbalance in Generative Listwise Recommendation via Discrepancy-Guided Exploration
Abstract
Generative listwise recommender systems search over a combinatorial space of item lists, yet sparse interaction feedback often leads to highly imbalanced exploration: training increasingly concentrates on well-covered lists, while potentially high-utility lists remain rarely sampled and insufficiently evaluated. We identify this sampling imbalance as a key bottleneck and show that the discrepancy between a list's empirical reward and its model-assigned selection probability provides a direct signal of under-exploration. Building on this insight, we propose a Teacher–Student collaborative framework (Tesco) for discrepancy-guided exploration. An auxiliary Teacher combines the Student's discrepancy signal with list reward to preferentially sample under-explored yet promising lists, while the Student learns from both self-generated and Teacher-generated experiences through alternating sampling with a shared replay buffer. This mechanism enables targeted exploration beyond the Student's current sampling distribution without modifying the original reward definition or observed interaction data. Experiments on three real-world recommendation datasets under both simulated online and offline settings show that Tesco consistently improves list-level utility and ranking quality while maintaining competitive diversity and coverage, with average-reward gains of up to 4.5% over the strongest baseline. Extensive ablation and sensitivity analyses further validate the effectiveness and robustness of the proposed discrepancy-guided exploration mechanism. The code is available at https://anonymous.4open.science/r/Tesco-77DD/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.