JOINT-EMBEDDING PREDICTIVE LEARNING CLOSES THE UNIMODAL–MULTIMODAL GAP IN ECG PRE- TRAINING
Abstract
In time series representation learning, pairing signals with free-text descriptions typically outperforms signal-only pretraining; in ECG, where most recordings carry a written interpretive report, the gap is read as evidence that language supplies information the waveform lacks. Whether it survives a controlled comparison is unknown: published comparisons vary the pretraining corpus, the evaluation protocol and whether baselines are retrained at all. We fix all three, retraining six published methods from their official code on 1.21 million recordings and probing every model with the same frozen linear probe on PTB-XL and SPH. Under these controls, joint-embedding predictive learning (JEPA) largely closes the gap: it comes within 0.005 macro-AUROC of the strongest multimodal model on PTB-XL and scores 0.007 above it on SPH, the two being equivalent under two one-sided tests with a 0.02 margin. Reconstruction, discrete-target and contrastive baselines trail behind. Adding reports gives no gain with all labels, but adds 0.015 and 0.045 with 10% of them. Language buys label efficiency rather than a higher ceiling: signal-only latent prediction reaches the same performance without ever seeing a report, and whether text earns its cost elsewhere in time series needs the same controls applied.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.