Detecting cardiac states with wearable photoplethysmograms and implications for out-of-hospital cardiac arrest detection
Bibliographic record
Abstract
Out-of-hospital cardiac arrest (OHCA) is a global health problem affecting approximately 4.4 million individuals yearly. OHCA has a poor survival rate, specifically when unwitnessed (accounting for up to 75% of cases). Rapid recognition can significantly improve OHCA survival, and consumer wearables with continuous cardiopulmonary monitoring capabilities hold potential to "witness" cardiac arrest and activate emergency services. In this study, we used an arterial occlusion model to simulate cardiac arrest and investigated the ability of infrared photoplethysmogram (PPG) sensors, often utilized in consumer wearable devices, to differentiate normal cardiac pulsation, pulseless cardiac (i.e., resembling a cardiac arrest), and non-physiologic (i.e., off-body) states. Across the classification models trained and evaluated on three anatomical locations, higher classification performances were observed on the finger (macro average F1-score of 0.964 on the fingertip and 0.954 on the finger base) compared to the wrist (macro average F1-score of 0.837). The wrist-based classification model, which was trained and evaluated using all PPG measurements, including both high- and low-quality recordings, achieved a macro average precision and recall of 0.922 and 0.800, respectively. This wrist-based model, which represents the most common form factor in consumer wearables, could only capture about 43.8% of pulseless events. However, models trained and tested exclusively on high-quality recordings achieved higher classification outcomes (macro average F1-score of 0.975 on the fingertip, 0.973 on the finger base, and 0.934 on the wrist). The fingertip model had the highest performance to differentiate arterial occlusion pulselessness from normal cardiac pulsation and off-body measurements with macro average precision and recall of 0.978 and 0.972, respectively. This model was able to identify 93.7% of pulseless states (i.e., resembling a cardiac arrest event), with a 0.4% false positive rate. All classification models relied on a combination of time-, power spectral density (PSD)-, and frequency-domain features to differentiate normal cardiac pulsation, pulseless cardiac, and off-body PPG recordings. However, our best model represented an idealized detection condition, relying on ensuring high-quality PPG data for training and evaluation of machine learning algorithms. While 90.7% of our PPG recordings from the fingertip were considered of high quality, only 53.2% of the measurements from the wrist passed the quality criteria. Our findings have implications for adapting consumer wearables to provide OHCA detection, involving advancements in hardware and software to ensure high-quality measurements in real-world settings, as well as development of wearables with form factors that enable high-quality PPG data acquisition more consistently. Given these improvements, we demonstrate that OHCA detection can feasibly be made available to anyone using PPG-based consumer wearables.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".