One Representation for Detection and Interpretation: Language-anchored Unified ECG Representation Learning via Signal, Image, and Clinical Text
Abstract
Electrocardiogram (ECG) representation is a core for computer-aided diagnosis of cardiovascular diseases (CVDs). The existing ECG-Text alignment-based methods are a promising paradigm for ECG representation learning. However, the learned representations struggle to concurrently execute both disease detection and interpretation in the diagnosis of CVDs, as proficiently as cardiologists do. Thus, we propose a LaUER, a Language-anchored Unified ECG Representation Learning framework for above tasks, like cardiologists. To approach the view of cardiologists, the ECG image is introduced to complement the ECG signal. Using clinical text, we design a multi-granularity language-anchored multimodal pre-training strategy to guide ECG representation learning in the language space, including a semantic-granularity learning, a concept-granularity learning, and a token-granularity learning. Through progressive guidance, a unified ECG representation is ultimately obtained that exhibits both detectability and interpretability. To evaluate the performance of LaUER, we conduct linear probing, zero-shot classification, and report generation evaluation on three ECG benchmark datasets, comparing it with the state-of-the-arts (SOTAs). Experimental results show that LaUER outperforms the SOTAs in both linear probing and zero-shot classification, while achieving performance close to SOTAs in report generation at a significantly lower cost compared to SOTAs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.