Medical Correctness and User Friendliness of Available Apps for Cardiopulmonary Resuscitation: Systematic Search Combined With Guideline Adherence and Usability Evaluation
Bibliographic record
Abstract
BACKGROUND: In case of a cardiac arrest, start of cardiopulmonary resuscitation by a bystander before the arrival of the emergency personnel increases the probability of survival. However, the steps of high-quality resuscitation are not known by every bystander or might be forgotten in this complex and time-critical situation. Mobile phone apps offering real-time step-by-step instructions might be a valuable source of information. OBJECTIVE: The aim of this study was to examine mobile phone apps offering real-time instructions in German or English in case of a cardiac arrest, to evaluate their adherence to current resuscitation guidelines, and to test their usability. METHODS: Our 3-step approach combines a systematic review of currently available apps guiding a medical layperson through a resuscitation situation, an adherence testing to medical guidelines, and a usability evaluation of the determined apps. The systematic review followed an adapted preferred reporting items for systematic reviews and meta-analyses flow diagram, the guideline adherence was tested by applying a conformity checklist, and the usability was evaluated by a group of mobile phone frequent users and emergency physicians with the system usability scale (SUS) tool. RESULTS: The structured search in Google Play Store and Apple App Store resulted in 3890 hits. After removing redundant ones, 2640 hits were checked for fulfilling the inclusion criteria. As a result, 34 apps meeting all inclusion criteria were identified. These included apps were analyzed to determine medical accuracy as defined by the European Resuscitation Council's guidelines. Only 5 out of 34 apps (15%, 5/34) fulfilled all criteria chosen to determine guideline adherence. All other apps provided no or wrong information on at least one relevant topic. The usability of 3 apps was evaluated by 10 mobile phone frequent users and 9 emergency physicians. Of these 3 apps, solely the app "HELP Notfall" (median=87.5) was ranked with an SUS score above the published average of 68. This app was rated significantly superior to "HAMBURG SCHOCKT" (median=55; asymptotic Wilcoxon test: z=-3.63, P<.01, n=19) and "Mein DRK" (median=32.5; asymptotic Wilcoxon test: z=-3.83, P<.01, n=19). CONCLUSIONS: Implementing a systematic quality control for health-related apps should be enforced to ensure that all products provide medically accurate content and sufficient usability in complex situations. This is of exceptional importance for apps dealing with the treatment of life-threatening events such as cardiac arrest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".