BiomECG: Learning Transferable ECG Representations from Small Cohorts
Abstract
Self-supervised learning (SSL) has become a standard approach to ECG representation learning, reducing the need for large clinically annotated datasets. Foundation models have been proposed based on very large cohorts. However, they may not be suited to every target population, making lightweight SSL pipelines that work with smaller datasets beneficial. We propose biomECG, an ECG biometric authentication framework for SSL pretraining and downstream evaluation, training compact Siamese encoders to classify pairs of recordings as belonging to the same or different patients. We obtain results comparable to those of a state-of-the-art ECG foundation model developed from over 10 million clinically annotated ECGs, using approximately three orders of magnitude fewer patients for pretraining and over 100-fold fewer parameters. Crucially, frozen encoders combined through late fusion of ECG leads often outperform full fine-tuning at limited and intermediate downstream training sizes. Late fusion also improves mortality discrimination over early fusion, particularly with smaller training cohorts. We evaluate in detail the impact of both the SSL pretraining dataset size and the downstream training dataset size using MIMIC and CODE-15. Pretraining on separate cohorts of similar size yields comparable transfer performance on MIMIC and CODE-15, indicating that representations learned through biometric pretraining capture information relevant across datasets. Code for biomECG, including pretraining and evaluation on MIMIC and CODE-15, is available at https://github.com/anonymoussubmission1000/biomECG-iclr-anon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.