Estimating the Risk of Lower Extremity Complications in Adults Newly Diagnosed With Diabetic Polyneuropathy: Retrospective Cohort Study
Bibliographic record
Abstract
Background: Diabetes-related lower extremity complications, such as foot ulceration and amputation, are on the rise, currently affecting nearly 131 million people worldwide. Methods for early detection of individuals at high risk remain elusive. While data-driven diabetic polyneuropathy algorithms exist, high-performing, clinically useful tools to assess risk are needed to improve clinical care. Objective: This study aimed to develop an electronic medical record-based machine learning algorithm that would predict lower extremity complications. Methods: We conducted a retrospective longitudinal cohort study to predict the risk of lower extremity complications within 24 months of an initial diagnosis of diabetic polyneuropathy. From an initial cohort of 468,162 individuals with at least 1 diagnosis of diabetic polyneuropathy at one of 2 multispecialty health care systems (based in northern California and Colorado) between April 2012 and December 2016, we created an analytic cohort of 48,209 adults with continuous enrollment, who were newly diagnosed with no evidence of end-of-life care. The outcome was any lower extremity complication, including foot ulceration, osteomyelitis, gangrene, or lower extremity amputation. We randomly split the data into training (38,569/48209; 80%) and testing (9,640/48209; 20%) datasets. In the training dataset, we used super Learner (SL), an ensemble learning method that employs cross-validation and combines multiple candidate risk predictors, into a single risk predictor. We evaluated the performance of the SL risk predictor in the testing dataset using the receiver operating characteristic curve and a calibration plot. Results: Of the 48,209 individuals in the cohort, 2327 developed a lower extremity complication during follow-up. The SL risk estimator exhibited good discrimination (AUC=0.845, 95% CI 0.826-0.863) and calibration. A modified version of our SL algorithm, simplified to facilitate real-world adoption, had only slightly reduced discrimination (AUC=0.817, 95%CI 0.797-0.837). The modified version slightly outperformed the naïve logistic regression model (AUC=0.804, 95% CI 0.783-0.825) in terms of precision gained relative to the frequency of alerts and number of patients that needed to be evaluated. Conclusions: We have built a machine learning-based risk estimator with the potential to improve clinical detection of diabetic patients at high risk for lower extremity complications at the time of an initial diabetic polyneuropathy diagnosis. The algorithm exhibited good discriminant validity and calibration using only data from the electronic medical record. Additional research will be needed to identify optimal contexts and strategies for maximizing algorithmic fairness in both interpretation and deployment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".