Comparison of 30-Day Mortality Models for Profiling Hospital Performance in Acute Ischemic Stroke With vs Without Adjustment for Stroke Severity
Bibliographic record
Abstract
CONTEXT: There is increasing interest in reporting risk-standardized outcomes for Medicare beneficiaries hospitalized with acute ischemic stroke, but whether it is necessary to include adjustment for initial stroke severity has not been well studied. OBJECTIVE: To evaluate the degree to which hospital outcome ratings and potential eligibility for financial incentives are altered after including initial stroke severity in a claims-based risk model for hospital 30-day mortality for acute ischemic stroke. DESIGN, SETTING, AND PATIENTS: Data were analyzed from 782 Get With The Guidelines-Stroke participating hospitals on 127,950 fee-for-service Medicare beneficiaries with ischemic stroke who had a score documented for the National Institutes of Health Stroke Scale (NIHSS, a 15-item neurological examination scale with scores from 0 to 42, with higher scores indicating more severe stroke) between April 2003 and December 2009. Performance of claims-based hospital mortality risk models with and without inclusion of NIHSS scores for 30-day mortality was evaluated and hospital rankings from both models were compared. MAIN OUTCOMES MEASURES: Model discrimination, hospital 30-day mortality outcome rankings, and value-based purchasing financial incentive categories. RESULTS: Across the study population, the mean (SD) NIHSS score was 8.23 (8.11) (median, 5; interquartile range, 2-12). There were 18,186 deaths (14.5%) within the first 30 days, including 7430 deaths (5.8%) during the index hospitalization. The hospital mortality model with NIHSS scores had significantly better discrimination than the model without (C statistic, 0.864; 95% CI, 0.861-0.867, vs 0.772; 95% CI, 0.769-0.776; P < .001). Among hospitals ranked in the top 20% or bottom 20% of performers by the claims model without NIHSS scores, 26.3% were ranked differently by the model with NIHSS scores. Of hospitals initially classified as having "worse than expected" mortality, 57.7% were reclassified to "as expected" by the model with NIHSS scores. The net reclassification improvement (93.1%; 95% CI, 91.6%-94.6%; P < .001) and integrated discrimination improvement (15.0%; 95% CI, 14.6%-15.3%; P < .001) indexes both demonstrated significant enhancement of model performance after the addition of NIHSS. Explained variance and model calibration was also improved with the addition of NIHSS scores. CONCLUSION: Adding stroke severity as measured by the NIHSS to a hospital 30-day risk model based on claims data for Medicare beneficiaries with acute ischemic stroke was associated with considerably improved model discrimination and change in mortality performance rankings for a substantial portion of hospitals.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".