SOLAR: SVD-Optimized Lifelong Attention for Recommendation
Abstract
Attention mechanism remains the defining operator in Transformers since it provides expressive global credit assignment, yet its quadratic cost in sequence length makes long-context modeling expensive and often forces truncation or other heuristics. Linear attention reduces complexity to by reordering computation through kernel feature maps, but this reformulation drops the softmax mechanism and shifts the attention score distribution. Lifelong recommendation requires efficient attention as well, for large-scale sequence modeling with user histories and candidate items under tight latency and resource constraints. We introduce SVD-Attention, a novel attention mechanism, and SOLAR, a set-aware framework built on it for lifelong recommendation. SVD-Attention factorizes low-rank embeddings into principal components, computes candidate-to-interest scores in compact space, and applies softmax over those scores. Its bilinear reduction is exact on the rank- reconstruction, while the normalized output approximates token-level softmax with an explicitly bounded residual. The resulting computation reduces the cost from to , and supports 12,000 behaviors and 3,000 candidates per request without filtering. SOLAR achieves the best among compared methods on RecFlow and MIND, delivers 0.8531 AUC at approximately 95th-percentile latency in an industrial evaluation, and yields business gains with a relative lift in Video Views in the real-world online A/B test.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.