BenchECG and xECG: A Standardized Benchmark and Baseline for ECG Foundation Models
Synopsis
This work introduces BenchECG, a standardized benchmark spanning eight public ECG datasets, 421,171 patients, 1,674,704 recordings, and ten tasks (classification, segmentation, detection, regression, survival analysis), used to evaluate public foundation models such as ST-MEM, ECG-JEPA, and ECGFounder; it also proposes xECG, a bidirectional xLSTM model pretrained with SimDINOv2 self-supervised learning, which achieves the best BenchECG score of 0.868±0.0030 (mean rank 1.50 under finetuning and 1.20 under linear probing), is the only public model to perform strongly across all datasets and task types, and leads on long-context tasks (sleep apnea AUROC 0.932±0.014; MIT-BIH arrhythmia F1 0.677±0.025) and computational efficiency (about 10x less time and about 7x less memory on PTB-XL).
Interpretation
Introduces BenchECG, the first comprehensive standardized benchmark for ECG foundation models, covering eight datasets, ten tasks, and diverse signal types. Whereas prior work used narrow task selections and inconsistent datasets, BenchECG unifies signal characteristics (1-12 leads, 100-500 Hz, seconds to hours), populations (Europe, USA, China, Brazil; healthy to ICU patients), and task types (classification, segmentation, detection, regression, survival analysis), and requires that evaluated models not be pretrained on any evaluation dataset. The benchmark comprises 8 public datasets, 421,171 patients, and 1,674,704 recordings; each model is trained 5 times with mean±SD reported, and pairwise differences are assessed with two-sided Welch t-tests (p<0.05).
Proposes xECG, an ECG foundation model based on a bidirectional xLSTM architecture pretrained with SimDINOv2 self-supervised learning. Unlike prior transformer-centric self-supervised ECG models, xECG uses alternating sLSTM/mLSTM blocks (s,s,m,m,s,s,m,m,s) to process temporal patches bidirectionally, and adapts SimDINOv2's teacher-student self-distillation (with patch-level, sample-level, and coding-rate regularization losses) to the ECG time-series domain. Pretraining used CODE, Chapman & Ningbo, and INCART (about 8 million ECGs) for 100 epochs with batch size 512; it achieves the highest BenchECG score of 0.868±0.0030.
xECG clearly outperforms transformer/CNN baselines on long-context tasks and shows stronger feature transferability under linear probing. On sleep apnea, xECG reaches AUROC 0.932±0.014, significantly above ST-MEM (0.702±0.020) and supervised xLSTM (0.853±0.022); on MIT-BIH arrhythmia under linear probing, F1 is 0.674±0.013 versus 0.436±0.036 for ST-MEM. These differences are reported with p-values (e.g., p=0.0000001, p=0.000032) based on mean±SD across 5 independent runs.
xECG is more computationally efficient than transformer baselines, and code and weights are released. On the same hardware (a single Nvidia L40S), xECG required about 5.1 hours of finetuning across the benchmark versus 26.1 hours for ST-MEM; on PTB-XL at the same number of training steps, it used about 10x less time (7 min vs 67 min) and about 7x less memory (4.2 GB vs 28.4 GB). Efficiency comparisons were made on unified hardware and identical training steps (30 epochs, batch size 96, 5430 steps), with parameter counts reported (xECG 57.0M, ST-MEM 85.2M, etc.).
Perspective
BenchECG targets researchers and engineering teams who want to compare ECG foundation models under unified conditions, and applies to evaluation scenarios coverable by public datasets; xECG's strengths center on long-context and multi-task generalization, with pretraining corpora of CODE, INCART, and Chapman & Ningbo. The work provides a reusable evaluation framework and baseline for subsequent model development and selection on longer signals, more modalities, and broader populations.
Pretraining data scales differ substantially across models (ECGFounder about 10M ECGs, xECG about 8M, ST-MEM and ECG-JEPA about 400,000), so it is difficult to determine which pretraining methodology is optimal; evaluation focuses on technical metrics and does not directly assess clinical utility or safety, requiring prospective validation; a fixed benchmark risks targeted tuning and is constrained by public dataset availability, which may not capture the full variability of real-world ECG practice.
