acceptodds
Under review as a conference paper at ICLR 2027

Improve Discriminative Recommendation through Better Generative Pretraining

Abstract

Generative next-item prediction (NTP) pretraining improves click-through rate (CTR) prediction even when only the item embeddings are transferred. This demonstrates that pretraining yields effective item representations. We decompose embedding gradients and find that positive-target supervision is mostly important. Specifically, we discover that item frequency and attention structure make great effects on gradients. Low item frequency causes the inefficient supervision on embedding while self-attention structure makes the gradient direction sensitive to the most recent interaction. We therefore propose two modifications of NTP pretraining, Frequency Adjustment (FA) and shared-query Cross attention (Cross). First, FA strengthens the output gradients for infrequent targets to mitigate the supervision imbalance. Second, Cross uses a learned shared query to aggregate user history, reducing the model’s reliance on recent interactions. Both methods keep the next-item objective and the transfer pipeline. Experiments on public datasets show that joint method improves mean downstream AUC over standard NTP.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.