Scaling ECG-Language Models to 12 Million Training Examples
Abstract
ECG-language models (ELMs) show promise for automated ECG interpretation, but the effects of data scale and training strategies remain underexplored. We introduce Orah, an ELM combining a SigLIP2-adapted ECG encoder (SigLEP) with a 4 billion parameter large language model (LLM). Orah is trained on approximately 12 million examples across 12 datasets, the largest ELM training corpus to our knowledge. We develop a multi-stage training recipe spanning SigLEP pretraining, Orah pretraining, and supervised fine-tuning (SFT). We investigate the contributions of data scale and Orah pretraining across three benchmarks. Orah exceeds the strongest reported baselines in all metrics on ECG-QA-CoT and the ECG-R1 Benchmark. On the ECG-Reasoning Benchmark, adding approximately 3,000 task-specific examples from the opposing ECG source improves Initial Diagnosis Accuracy (IDA) by 27.26 points on PTB-XL and 30.45 points on MIMIC-IV-ECG, highlighting the value of small amounts of targeted supervision. Furthermore, our ablations show that increasing training data and computation under our recipe consistently improves performance across all three benchmarks. These results establish Orah as a strong ELM and provide empirical guidance for scaling data and training ELMs. Upon acceptance, we will release all publicly available datasets, model weights, and code for preprocessing, training, and evaluation used throughout this study.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.