Evaluating Outcome Prediction Models in Endovascular Stroke Treatment Using Baseline, Treatment, and Posttreatment Variables
Bibliographic record
Abstract
Background: Statistical models to predict outcomes after endovascular therapy for acute ischemic stroke often incorporate baseline (pretreatment) variables only. We assessed the performance of stroke outcome prediction models for endovascular therapy in stroke in an iterative fashion using baseline, treatment-related, and posttreatment variables. Methods: Data from the ESCAPE-NA1 (Safety and Efficacy of Nerinetide [NA-1] in Subjects Undergoing Endovascular Thrombectomy for Stroke) trial were used to build 4 outcome prediction models using multivariable logistic regression: model 1 included baseline variables available before treatment decision making, model 2 included additional treatment-related variables, model 3 additional posttreatment variables that become available early (within 24-48 hours), and model 4 later (beyond 48 hours) after endovascular therapy. The primary outcome was functional independence (90-day Modified Rankin Scale score 0-2). Model performance was compared using the area under the receiver operating characteristic curve (AUC). Shapley values were used to determine marginal contributions of variables to outcome variance in the regression models. Results: Among 1105 patients, functional independence was achieved by 666 (60.3%). When using baseline variables only (model 1), the AUC was 0.74 (95% CI, 0.71-0.77); this iteratively improved when treatment and posttreatment variables were added to the models (model 2: AUC, 0.77; 95% CI, 0.74-0.80; model 3: AUC, 0.80; 95% CI, 0.77-0.83; model 4: AUC, 0.82; 95% CI, 0.79-0.85). With baseline variables alone, 26% of patients who achieved functional independence were erroneously classified as not achieving functional independence. Even with the most comprehensive model, 19.8% of patients were misclassified as such. Patient age contributed most to outcome variance (Shapley value, 0.28), followed by severe adverse events including pneumonia (0.16) and intracranial hemorrhage at 24-hours imaging (0.13). Conclusions: A substantial contribution to outcomes after endovascular therapy comes from factors unrelated to currently collected baseline patient variables. One-fifth of patients achieving functional independence were misclassified as not achieving independence, even with the most comprehensive model. Our findings suggest that the achievable accuracy of current outcome prediction models is limited, and caution should be used when applying them in clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".