Generalize or Remember? Rethinking Generative Recommendation from a Data-Centric Perspective
Abstract
abstract Modern generative recommendation (GR) is largely built around generalization: semantic IDs (SIDs) allow related items to share discrete tokens and transfer knowledge across items. Yet generalization alone is insufficient for recommendation: when interaction patterns recur across users, effectively reusing such recurring knowledge is equally important. This raises a key question: when should a generative recommender generalize, and when should it remember? We revisit GR from a data-centric perspective by distinguishing interaction instances according to whether their target transitions are supported by recurring cross-user behaviors, yielding memory and generalization subsets. Our analysis shows that memory specialization improves performance on the memory subset at the cost of generalization and overall performance, while gradients from the two subsets frequently conflict. To address this trade-off, we propose MemOPSD, a memory-driven on-policy self-distillation framework. MemOPSD trains a memory-specialized teacher on recurring interactions and selectively transfers its knowledge only to memory samples while retaining the original GR objective on all training instances, strengthening memory without indiscriminately imposing memory-oriented supervision on generalization data. Since conventional distillation relies on fixed prefixes that differ from student-generated states at inference time, MemOPSD further performs on-policy distillation along student-generated SID trajectories, transferring memory knowledge at the states actually visited by the student. Extensive experiments on three benchmark datasets and four GR backbones demonstrate consistent improvements in overall and memory recommendation performance. The code is available at https://anonymous.4open.science/r/MemOPSD-0627. abstract
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.