SE-JEPA: Learning Representations of Structured Event Sequences via Latent Prediction
Abstract
Structured event data contains within-event field dependencies and cross-event contextual relationships that can provide complementary information for representation learning. A key challenge is to integrate information from related events into field-level predictive learning. We propose the Structured-Event Joint-Embedding Predictive Architecture (SE-JEPA), a self-supervised framework that combines local field modeling with cross-event context through masked-field latent prediction. SE-JEPA constructs within-event representations using field-specific encodings and retrieves field-aligned context through same-field cross-event attention. An event-level residual auxiliary objective further uses pooled event representations to learn context-induced corrections to local predictions. The two mechanisms optimize a shared encoder and predictor using latent targets supplied by an exponential moving-average teacher. Experiments on four public structured-event benchmarks covering diverse downstream tasks demonstrate the effectiveness of SE-JEPA under frozen linear evaluation and full fine-tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.