acceptodds
Under review as a conference paper at ICLR 2027

ECGQuest: Benchmarking and Fine-Tuning Language Models for Electrocardiography

Abstract

Electrocardiogram (ECG) interpretation requires knowledge of cardiology, electrophysiology, clinical diagnosis, ECG waveforms, signal acquisition, and ECG instrumentation. While large language models have proven effective for such applications, existing language-models primarily assess broad medical knowledge or the interpretation of individual ECG signals and images, rather than comprehensive context knowledge required for ECG interpretation. We built ECGQuest, a literature-grounded resource for evaluating and fine-tuning ECG-specific language models. An automated GPT-4o-based pipeline generated questions from 23 ECG references and proceedings of the Computing in Cardiology conference between 2003–2025. The final dataset contains 10,904 unique True/False questions paired with their negated form derived from the same source (21,808 Q&A in total). We evaluated three commercial and 20 open-source language models on the held-out ECGQuest test set in a zero-shot setting. We also fine-tuned five open-source models with 7–14 B parameters using Low-Rank Adaptation and included BERT and BiomedBERT as supervised encoder baselines. Generalization was evaluated using ECG-related subsets of MedMCQA and MedQA datasets converted to binary True/False questions using their official examination answer keys. Zero-shot accuracy on ECGQuest ranged from 49.5% to 74.4%, with GPT-5 achieving the highest accuracy. General-purpose models outperformed all medically specialized models, and several models showed strong directional True/False biases. Encoder baselines performed near chance. Fine-tuning improved all open-source models on ECGQuest by 6.5–14.1%. The fine-tuned DeepSeek-R1-Distill-Qwen-14 B model achieved 76.3% accuracy, and a five-model voting ensemble achieved the highest overall accuracy of 78.5%. On MedMCQA and MedQA, fine-tuning primarily benefited weak or class-biased models and did not consistently improve already strong base models. ECGQuest is a reproducible benchmark for evaluating contextual ECG knowledge in language models and demonstrates that targeted, parameter-efficient fine-tuning can make small language models competitive with substantially larger commercial models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.