Machine Learning to Predict Outcomes of Endovascular Intervention for Patients With PAD
Bibliographic record
Abstract
Importance: Endovascular intervention for peripheral artery disease (PAD) carries nonnegligible perioperative risks; however, outcome prediction tools are limited. Objective: To develop machine learning (ML) algorithms that can predict outcomes following endovascular intervention for PAD. Design, Setting, and Participants: This prognostic study included patients who underwent endovascular intervention for PAD between January 1, 2004, and July 5, 2023, with 1 year of follow-up. Data were obtained from the Vascular Quality Initiative (VQI), a multicenter registry containing data from vascular surgeons and interventionalists at more than 1000 academic and community hospitals. From an initial cohort of 262 242 patients, 26 565 were excluded due to treatment for acute limb ischemia (n = 14 642) or aneurysmal disease (n = 3456), unreported symptom status (n = 4401) or procedure type (n = 2319), or concurrent bypass (n = 1747). Data were split into training (70%) and test (30%) sets. Exposures: A total of 112 predictive features (75 preoperative [demographic and clinical], 24 intraoperative [procedural], and 13 postoperative [in-hospital course and complications]) from the index hospitalization were identified. Main Outcomes and Measures: Using 10-fold cross-validation, 6 ML models were trained using preoperative features to predict 1-year major adverse limb event (MALE; composite of thrombectomy or thrombolysis, surgical reintervention, or major amputation) or death. The primary model evaluation metric was area under the receiver operating characteristic curve (AUROC). After selecting the best performing algorithm, additional models were built using intraoperative and postoperative data. Results: Overall, 235 677 patients who underwent endovascular intervention for PAD were included (mean [SD] age, 68.4 [11.1] years; 94 979 [40.3%] female) and 71 683 (30.4%) developed 1-year MALE or death. The best preoperative prediction model was extreme gradient boosting (XGBoost), achieving the following performance metrics: AUROC, 0.94 (95% CI, 0.93-0.95); accuracy, 0.86 (95% CI, 0.85-0.87); sensitivity, 0.87; specificity, 0.85; positive predictive value, 0.85; and negative predictive value, 0.87. In comparison, logistic regression had an AUROC of 0.67 (95% CI, 0.65-0.69). The XGBoost model maintained excellent performance at the intraoperative and postoperative stages, with AUROCs of 0.94 (95% CI, 0.93-0.95) and 0.98 (95% CI, 0.97-0.99), respectively. Conclusions and Relevance: In this prognostic study, ML models were developed that accurately predicted outcomes following endovascular intervention for PAD, which performed better than logistic regression. These algorithms have potential for important utility in guiding perioperative risk-mitigation strategies to prevent adverse outcomes following endovascular intervention for PAD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".