acceptodds
Under review as a conference paper at ICLR 2027

Scaling ECG-Language Models to 12 Million Training Examples

Abstract

ECG-language models (ELMs) show promise for automated ECG interpretation, but the effects of data scale and training strategies remain underexplored. We introduce Orah, an ELM combining a SigLIP2-adapted ECG encoder (SigLEP) with a 4 billion parameter large language model (LLM). Orah is trained on approximately 12 million examples across 12 datasets, the largest ELM training corpus to our knowledge. We develop a multi-stage training recipe spanning SigLEP pretraining, Orah pretraining, and supervised fine-tuning (SFT). We investigate the contributions of data scale and Orah pretraining across three benchmarks. Orah exceeds the strongest reported baselines in all metrics on ECG-QA-CoT and the ECG-R1 Benchmark. On the ECG-Reasoning Benchmark, adding approximately 3,000 task-specific examples from the opposing ECG source improves Initial Diagnosis Accuracy (IDA) by 27.26 points on PTB-XL and 30.45 points on MIMIC-IV-ECG, highlighting the value of small amounts of targeted supervision. Furthermore, our ablations show that increasing training data and computation under our recipe consistently improves performance across all three benchmarks. These results establish Orah as a strong ELM and provide empirical guidance for scaling data and training ELMs. Upon acceptance, we will release all publicly available datasets, model weights, and code for preprocessing, training, and evaluation used throughout this study.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.