Skip to main content
Back to timeline
arXivSource publication:

MS-ECG-FM aligns ECG to multiple clinical report types and lifts macro-averaged AUROC to 93.6 across 98 detection tasks

Synopsis

The authors introduce MS-ECG-FM, an ECG foundation model pre-trained by contrastive alignment to several clinical note types — ECG machine reports, cardiologist ECG reports, echocardiography, chest X-ray and discharge summaries; on linear-probe evaluation over 98 labels across multiple datasets it reaches a macro-averaged AUROC of 93.6 versus 91.5 for ECGFounder and 91.2 for MELP, with clear gains on structural heart disease and maintained leadership in single-lead and reduced-lead configurations.

Source-provided article image: MS-ECG-FM: Towards a More Universal Electrocardiogram Foundation Model for Health Monitoring using Multi-source Contrastive Learning
Figure 1 ·

Figure 1: Conceptual overview of MS-ECG-FM. MS-ECG-FM is an ECG foundation model that can support a broader range of ECG detection tasks. It is pre-trained using concurrent alignment to ECG, ECHO, chest X-ray and discharge reports to improve the universality of its representations.

arXiv

Interpretation

Multi-source contrastive alignment lets a single ECG encoder absorb information from several clinical report types, reaching a macro-averaged AUROC of 93.6 over the 98-label evaluation set, above ECGFounder (91.5) and MELP (91.2), which uses the same pre-training data. Prior ECG foundation models relied exclusively on machine-generated ECG interpretation reports as their sole supervision; this work aligns the waveform to ECG, cardiologist ECG, echocardiography, chest X-ray and discharge reports jointly. Evaluated by linear probing on frozen embeddings across held-out datasets including PTB-XL, CPSC2018, CSN and EchoNext-Mini, with paired patient-level bootstrap significance testing; MS-ECG-FM pre-trains on only 720,102 ECGs versus the 800,035 used by the compared methods.

Different report types favor different diagnostic domains: ECG cardiologist reports (AUROC 89.7) and ECG machine reports (89.5) are strongest for arrhythmias, conduction disturbances, ischemia/ST-T changes and voltage-criteria conditions, while discharge summaries and chest X-ray reports are most useful for structural heart disease and the exploratory MIMIC-EXP outcomes. The work turns the general claim that clinical text supervision matters into a per-label, per-report-type controlled comparison, and shows the multi-source model approaches the per-label oracle of the best single source. Single-source models share MS-ECG-FM's architecture and recipe and differ only in the aligned report type, forming a controlled comparison; report availability differs substantially (ECHO 81,821 notes from 34,147 patients versus 331,793 discharge summaries).

MS-ECG-FM retains its lead in reduced-lead settings, with 12-lead macro-averaged AUROC 93.6 and single-lead I, II and V2 at 86.9, 88.1 and 85.0, all above the strongest baselines. The tokenizer produces lead-level tokens with weights shared across leads, and pre-training uses random lead masking, so representations learned from 12-lead data transfer to arbitrary reduced-lead subsets. Uses the PhysioNet/CinC Challenge 2021 single-lead and reduced-lead subset configurations, reported by diagnostic group, alongside per-label AUROC and AUPRC.

Even when aligned only to ECG machine reports, MS-ECG-FM's architecture and recipe reach a 98-label macro-averaged AUROC of 92.8, above MERL (88.8), D-BETA (90.3) and MELP (91.2) pre-trained on the same machine reports. This isolates contributions beyond multi-source alignment: frozen MedGemma-27B text embeddings, no input normalization so amplitude information is preserved for voltage-criteria diagnoses, and a different backbone and pre-training recipe. Ablations show that using an online text encoder or a zero-shot checkpoint selection criterion degrades performance noticeably, while a frozen text encoder with RankMe checkpoint selection gives the strongest representations.

Perspective

The results apply to research and modeling settings that use clinical 12-lead ECGs with accompanying clinical text, and to deployment scenarios that transfer 12-lead representations to single-lead or reduced-lead subsets such as bedside monitors, Holter monitors, patches or wrist-worn devices. The authors note that hospital systems already store multimodal notes at far greater scale than was available here, and that pre-training a larger MS-ECG-FM on such text could improve performance beyond the gains shown. The evaluation framework itself — diagnostic groupings, performance profiles and paired bootstrap testing — is offered as a shared tool for comparing ECG foundation models.

Pre-training data come from a single health system, skewed toward an older population (median age 58, interquartile range 43–72) and toward inpatient or emergency visits, so transfer to other populations and hospital systems may still be attenuated depending on the distribution shift. Reduced-lead evaluations were derived from clinical 12-lead recordings rather than natively acquired wearable or ambulatory ECGs, so validation on real-world reduced-lead data remains open. Some low-prevalence but clinically acute conditions are not adequately represented in the available evaluation datasets; the randomly initialized encoder was not significantly worse on right ventricular hypertrophy, hypertrophic cardiomyopathy and cardiac amyloidosis, which the authors attribute to very low prevalence. Foundation models are not inherently interpretable, and the authors note this is not the full solution to the adoption gap.

Sources