LIFT-JEPA: Task-Aware Joint-Embedding Supervision via Target Interventions for Spatiotemporal Forecasting
Abstract
In spatiotemporal forecasting where task-critical phenomena are sparse, prediction losses applied in observation space allow pixel-wise errors in task-critical regions to be directly reweighted. By contrast, joint-embedding predictive learning applies the loss in representation space; because latent units lack an explicit correspondence to task semantics expressed in observation space, observation-space task weights cannot be directly applied to the representation-matching loss. To this end, we introduce LIFT-JEPA (Latent Importance From Target Interventions for Joint-Embedding Prediction). Given a future observation, LIFT-JEPA constructs an intervened counterpart by removing task-critical structures and encodes both using the same target encoder. The resulting representation differences are used to estimate the task relevance of individual representation units and reweight their matching losses, thereby translating task priorities specified in observation space into representation-space supervision. We evaluate LIFT-JEPA across precipitation nowcasting and urban traffic flow prediction using three forecasting backbones with distinct inductive biases. Compared with Vanilla JEPA and alternative task-weighting strategies, LIFT-JEPA consistently improves task-relevant forecasting performance across both domains, from intensity-defined targets in radar nowcasting to temporal-drift targets in traffic flow prediction. These results demonstrate that LIFT-JEPA can translate task semantics ranging from spatial intensity structures to temporal drift into effective representation-space supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.