Ability of Current Machine Learning Algorithms to Predict and Detect Hypoglycemia in Patients With Diabetes Mellitus: Meta-analysis
Bibliographic record
Abstract
BACKGROUND: Machine learning (ML) algorithms have been widely introduced to diabetes research including those for the identification of hypoglycemia. OBJECTIVE: The objective of this meta-analysis is to assess the current ability of ML algorithms to detect hypoglycemia (ie, alert to hypoglycemia coinciding with its symptoms) or predict hypoglycemia (ie, alert to hypoglycemia before its symptoms have occurred). METHODS: Electronic literature searches (from January 1, 1950, to September 14, 2020) were conducted using the Dialog platform that covers 96 databases of peer-reviewed literature. Included studies had to train the ML algorithm in order to build a model to detect or predict hypoglycemia and test its performance. The set of 2 × 2 data (ie, number of true positives, false positives, true negatives, and false negatives) was pooled with a hierarchical summary receiver operating characteristic model. RESULTS: A total of 33 studies (14 studies for detecting hypoglycemia and 19 studies for predicting hypoglycemia) were eligible. For detection of hypoglycemia, pooled estimates (95% CI) of sensitivity, specificity, positive likelihood ratio (PLR), and negative likelihood ratio (NLR) were 0.79 (0.75-0.83), 0.80 (0.64-0.91), 8.05 (4.79-13.51), and 0.18 (0.12-0.27), respectively. For prediction of hypoglycemia, pooled estimates (95% CI) were 0.80 (0.72-0.86) for sensitivity, 0.92 (0.87-0.96) for specificity, 10.42 (5.82-18.65) for PLR, and 0.22 (0.15-0.31) for NLR. CONCLUSIONS: Current ML algorithms have insufficient ability to detect ongoing hypoglycemia and considerate ability to predict impeding hypoglycemia in patients with diabetes mellitus using hypoglycemic drugs with regard to diagnostic tests in accordance with the Users' Guide to Medical Literature (PLR should be ≥5 and NLR should be ≤0.2 for moderate reliability). However, it should be emphasized that the clinical applicability of these ML algorithms should be evaluated according to patients' risk profiles such as for hypoglycemia and its associated complications (eg, arrhythmia, neuroglycopenia) as well as the average ability of the ML algorithms. Continued research is required to develop more accurate ML algorithms than those that currently exist and to enhance the feasibility of applying ML in clinical settings. TRIAL REGISTRATION: PROSPERO International Prospective Register of Systematic Reviews CRD42020163682; http://www.crd.york.ac.uk/PROSPERO/display_record.php?ID=CRD42020163682.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.056 | 0.127 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.013 | 0.060 |
| Bibliometrics | 0.008 | 0.006 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".