Defining the Limits of Pre-Transplant Risk Prediction in AML: Evidence From Machine Learning and Regression Models
Bibliographic record
Abstract
BACKGROUND: Allogeneic hematopoietic cell transplantation (allo-HCT) for acute myeloid leukemia (AML) is associated with considerable morbidity and mortality. Machine learning (ML) techniques are increasingly applied to predict outcomes in medicine. OBJECTIVES: To evaluate the role of ML in predicting overall survival (OS) after allo-HCT for AML and compare ML with traditional Cox regression. STUDY DESIGN: Using an internal cohort of 2253 patients and 14 pre-allo-HCT variables, we developed three models: Cox regression with time-varying coefficients (Cox-TVC), Elastic-net Cox Regression (Cox-EN), and Random Survival Forest (RSF). Performance was evaluated using multiple metrics including C-index, net reclassification improvement (NRI) and decision curve analysis (DCA). Patients were stratified into tertiles of model predicted 24-month mortality. External validation was performed in 252 single-center patients with uniform measurable residual disease (MRD) assessment. RESULTS: Model-derived risk scores strongly correlated (r = 0.886 to 0.963). Across models, age ≥ 60 years, MRD positivity, and adapted European LeukemiaNet (aELN) adverse risk were the strongest predictors. Effects of age attenuated over time (HR for age ≥60: 2.59 [95% CI: 1.54 to 4.35] at 1 year, 1.92 [95% CI: 1.04 to 3.56] at 5 years). Compared with Hematopoietic Cell Transplant Comorbidity-Index (HCT-CI) and aELN, models improved risk stratification (NRI: 31% to 44%, p < .001). Discrimination remained modest but was higher in external cohort compared with internal cohort (0.69 to 0.71 versus 0.60 to 0.61), likely reflecting uniform MRD assessment. At a 25%, risk threshold for clinical decision-making, models identified approximately 1 additional high-risk patient per 100 versus HCT-CI, or aELN. At 2-years, 25% to 27% of patients categorized as low-risk had died, while 45% to 48% categorized as high-risk were alive. CONCLUSIONS: ML approaches improved risk stratification over HCT-CI and aELN but performed comparably with Cox model. Individual outcome prediction using static pre-transplant models remained modest. Progress will require MRD standardization, richer data, and dynamic peri- and post-transplant modeling.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".