WORD-CTR: Write Once, Read Deep for Efficient Long-Sequence Click-Through Rate Prediction
Abstract
Click-through rate (CTR) prediction often scores many candidate items for the same user, making repeated access to long behavior histories costly. Reusing a candidate-independent user state can reduce this repeated work, provided that different candidates can still obtain distinct information from it. We present WordCTR (Write Once, Read Deep), which represents the retained behavior history as a fixed-size, multi-head associative memory. Its writer fits a linear key-to-value map in each head using content- and recency-weighted ridge regression. Stacked reader blocks combine non-sequential user, candidate, and context fields and make successive queries to the same memory without revisiting the behavior sequence. For fixed model dimensions, candidate-side reading is independent of history length once the memory is available, while memory construction remains history-dependent. We evaluate WordCTR against 17 baselines on three public datasets and an InHouse dataset that includes a behavior stream of up to 1024 events. WordCTR obtains the best reported AUC, GAUC, and NE on the public datasets, and the best AUC and NE on InHouse. We also examine cached and uncached inference across history lengths and candidate counts. In the measured single-request GPU workload, candidate-side read-and-score time varies little across tested history lengths, while caching reduces total latency with a larger benefit at higher tested candidate counts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.