acceptodds
Under review as a conference paper at ICLR 2027

Evolving Supervision through Learning Feedback for Generative Recommendation

Abstract

Generative recommendation models commonly learn target-item distributions through next-token prediction (NTP) on historical interactions. Yet observed interactions are incomplete proxies for users' latent preferences, rather than ground-truth descriptions of their preferences. Supervision can be enriched by extending click-based targets to purchases or incorporating behavioral signals from other contexts. The usefulness of additional supervision depends both on what information it adds and on what the current model has already learned. This motivates us to jointly consider supervision construction and feedback-guided learning. We propose a feedback-driven supervision evolution framework with two nested cycles. The outer Supervision Evolution Cycle (SEC) uses large language models (LLMs) to analyze updated performance feedback, diagnose supervision needs, and guide subsequent supervision construction. Within each round, the inner Feedback-Guided Learning Cycle (FGLC) updates the policy for incorporating constructed supervision based on its observed learning effects. We apply the framework to expand supervision in a recommendation retrieval system. Multi-round offline experiments on industrial recommendation retrieval show that the framework improves retrieval performance under matched total added supervision. Online A/B testing also shows improvements in ad clicks and revenue. The system has been fully deployed on a large-scale industrial e-commerce platform serving hundreds of millions of active users.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.