acceptodds
Under review as a conference paper at ICLR 2027

ELiT: An Expressive Lite Transformer for CTR Prediction with Error-Bounded Attention Approximation

Abstract

Recommender systems (RS) have long been dominated by the Embedding & MLP paradigm, which is now approaching its intrinsic performance ceiling. Motivated by the scaling laws of LLMs, recent studies have begun to adapt Transformer backbones to RS. However, direct transplantation incurs prohibitive parameter overhead, as heterogeneous input features rule out globally shared QKV projections and necessitate token-level parameterization, thereby causing parameter explosion, whereas lightweight variants typically sacrifice global interaction modeling. To resolve this trade-off, we propose the Expressive Lite Transformer (ELiT), a Transformer variant tailored to heterogeneous features that reconciles strong representational power. Our design is motivated by a key empirical insight: in RS, inter-token affinities are governed predominantly by feature fields rather than by specific feature values, rendering the inter-token affinity relationships largely invariant across different inputs and thus permitting aggressive pruning of redundant self-attention (SA) parameters. Concretely, ELiT introduces: (i) a progressive two-stage token mixing mechanism, fortified by a sparse MoE module, that substitutes for SA in modeling inter-token relations, thereby achieving sample-adaptive global interaction with negligible overhead; (ii) theoretical guarantees that the approximation error of the attention matrix remains tightly bounded, rigorously justifying its lightweight yet highly expressive architecture; (iii) a masked SwiGLU module that leverages the parameters freed by the simplified SA in (i) to upgrade the standard FFN for higher representational capacity. To avoid the substantial inference overhead of vanilla SwiGLU’s dual weight matrices, we replace them with a shared FP16 matrix and binary masks, alleviating the memory-bandwidth bottleneck at negligible computational cost. Extensive experiments on large-scale industrial and public datasets and online A/B tests consistently show that ELiT outperforms SOTA baselines, achieving an exceptional efficiency–expressiveness trade-off in real-world RS.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.