Improve Discriminative Recommendation through Better Generative Pretraining
Abstract
Generative next-item prediction (NTP) pretraining improves click-through rate (CTR) prediction even when only the item embeddings are transferred. This demonstrates that pretraining yields effective item representations. We decompose embedding gradients and find that positive-target supervision is mostly important. Specifically, we discover that item frequency and attention structure make great effects on gradients. Low item frequency causes the inefficient supervision on embedding while self-attention structure makes the gradient direction sensitive to the most recent interaction. We therefore propose two modifications of NTP pretraining, Frequency Adjustment (FA) and shared-query Cross attention (Cross). First, FA strengthens the output gradients for infrequent targets to mitigate the supervision imbalance. Second, Cross uses a learned shared query to aggregate user history, reducing the model’s reliance on recent interactions. Both methods keep the next-item objective and the transfer pipeline. Experiments on public datasets show that joint method improves mean downstream AUC over standard NTP.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.