SCALING TRANSFERABLE BEHAVIORAL REPRESENTATIONS: BRIDGE THE SEMANTIC GAP FROM NEXTEVENT PREDICTION TO DOWNSTREAM TASKS
Abstract
Internet platforms collect abundant behavioral data from a broad population, yet labels for downstream decision tasks are often scarce, delayed, and expensive. A central question is whether large-scale behavioral pre-training can learn representations that transfer across the substantial semantic gap between self-supervised behavior prediction and downstream objectives. We systematically study how pre-training objective, data scale, and model scale shape such cross-task transfer. We introduce BRIDGE (Behavioral Representations for Improving Downstream Generalization and Efficiency), which learns transferable representations of people from large-scale heterogeneous behavioral sequences together with structured features (static or slowly varying person-level attributes, commonly termed userprofile features in recommendation and credit-risk modeling) through autoregressive next-event prediction (NEP). We compare NEP with representative masked, multi-token, distribution-based, and embedding-based objectives. Objective preference depends on scale: embedding-based pre-training leads at the smaller evaluated scale, whereas NEP overtakes it at the larger scale. Despite the weak semantic alignment between NEP and downstream decision objectives, the learned representations transfer effectively to tasks such as credit-risk assessment and payment fraud detection. Across the evaluated scales, downstream out-of-time (OOT) performance improves consistently with both pre-training data volume and model size, while lower pre-training loss is associated with better downstream discrimination. With only 10% of the credit-risk labels, BRIDGE improves the OOT Kolmogorov–Smirnov (KS) statistic by 5.31% relative to end-to-end (E2E) supervised training under the same label budget and even surpasses E2E trained on the full labeled dataset. These results show that scaling behavioral pre-training can substantially improve cross-task transfer and label efficiency, while suggesting that the preferred pre-training objective itself can depend on scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.