Toward Robust Automated Cardiovascular Arrhythmia Detection using Self-supervised Learning and 1-Dimensional Vision Transformers
Bibliographic record
Abstract
Cardiovascular diseases are the primary cause of death globally. With the prevalence of electrocardiogram (ECG) machines within and outside the clinical environment, it is now possible to passively monitor a patient's heartbeat for cardiovascular diseases. The goal of this work is to emphasize the importance of selfsupervised learning for arrhythmia detection, leveraging the large amounts of unlabelled data recently made publicly available and demonstrating significant performance improvements as it reduces overfitting to class imbalance and noise. We propose Masked Patch Modelling (MPM) and leverage 8.2 million unlabelled ECGs to perform large-scale self-supervised pre-training and create a foundational 1dimensional Transformer model, PatchECG, that can be fine-tuned for any downstream tasks involving ECG data. We obtain state-of-the-art results on standard benchmark datasets, including PTB-XL multi-label classification, while setting new benchmarks on the largest and highest quality multi-label classification dataset to date. We find that PatchECG outperforms the current state-of-the-art with regard to computational efficiency, requiring only 1/5 of the computational resources while increasing model capacity by a factor of 14. We also compare the 1-dimensional PatchECG model to a state-of-the-art 2-dimensional vision Transformer and observe significantly higher performance. Finally, we perform ablation studies to investigate other methods for addressing the critical issues incurred with automated arrhythmia detection, resulting in a performance improvement of more than 2% under conditions of class imbalance, label noise, and over-parameterization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".