One Trajectory, Many Targets: Supervision Density and the Compute-Optimal Regime in Sequential Recommendation
Abstract
Sequential recommendation turns each user history into several next-item prediction tasks. Prefix-wise training encodes overlapping histories repeatedly, so grouping more targets into one update can save computation. The central question is how this supervision density changes the updates and compute needed to reach a fixed ranking quality. We identify an empirical scaling law and make its denser updates executable through exact shared computation. In five-seed MovieLens 1M (ML-1M), with the data, model, objective, initialization, and target-event budget fixed, quality trajectories align after updates are rescaled by density; threshold crossings follow an inverse-density relation after a shared offset of about 202 updates. At 0.966 million target events, m=2 reaches normalized discounted cumulative gain at rank 10 (NDCG@10) of 0.1519 in 102.8 training seconds, versus 0.1498 in 202.5 seconds for m=1. Fixed-state scans and five-seed SASRec sweeps show that the best density depends on the compute regime: the measured frontier peaks at m=4 (0.1201 NDCG@10) on ML-10M and improves through the tested m=16 (0.0463) on Amazon Beauty. *Exact Prefix-Loss Compilation* (EPC) evaluates aligned prefix losses in one causal pass while preserving a fixed update's loss and gradient when the states and training semantics match. Across 48 matched backbone–dataset comparisons, packing shortens train-to-best time; on Beauty, the combined configuration improves NDCG@10 by 52.8% and reaches its selected checkpoint 50.9× sooner. On ML-10M, four backbones reach matched quality in all five seeds with 1.30–4.56× fewer target events and 28.70–87.42× less wall time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.