ELiT: An Expressive Lite Transformer for CTR Prediction with Error-Bounded Attention Approximation
Abstract
Recommender systems (RS) have long been dominated by the Embedding & MLP paradigm, which is now approaching its intrinsic performance ceiling. Motivated by the scaling laws of LLMs, recent studies have begun to adapt Transformer backbones to RS. However, direct transplantation incurs prohibitive parameter overhead, as heterogeneous input features rule out globally shared QKV projections and necessitate token-level parameterization, thereby causing parameter explosion, whereas lightweight variants typically sacrifice global interaction modeling. To resolve this trade-off, we propose the Expressive Lite Transformer (ELiT), a Transformer variant tailored to heterogeneous features that reconciles strong representational power. Our design is motivated by a key empirical insight: in RS, inter-token affinities are governed predominantly by feature fields rather than by specific feature values, rendering the inter-token affinity relationships largely invariant across different inputs and thus permitting aggressive pruning of redundant self-attention (SA) parameters. Concretely, ELiT introduces: (i) a progressive two-stage token mixing mechanism, fortified by a sparse MoE module, that substitutes for SA in modeling inter-token relations, thereby achieving sample-adaptive global interaction with negligible overhead; (ii) theoretical guarantees that the approximation error of the attention matrix remains tightly bounded, rigorously justifying its lightweight yet highly expressive architecture; (iii) a masked SwiGLU module that leverages the parameters freed by the simplified SA in (i) to upgrade the standard FFN for higher representational capacity. To avoid the substantial inference overhead of vanilla SwiGLU’s dual weight matrices, we replace them with a shared FP16 matrix and binary masks, alleviating the memory-bandwidth bottleneck at negligible computational cost. Extensive experiments on large-scale industrial and public datasets and online A/B tests consistently show that ELiT outperforms SOTA baselines, achieving an exceptional efficiency–expressiveness trade-off in real-world RS.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.