Differential Impact of CD34+ Cell Dose for Different Age Groups in Allogeneic Hematopoietic Cell Transplantation for Acute Leukemia: A Machine Learning-Based Discovery
Bibliographic record
Abstract
Allogeneic hematopoietic cell transplantation (alloHCT) is curative for hematologic malignancies; however, it is associated with risks and complications. Identification of risk factors remain the topic of ongoing investigation. Specifically, CD34+ cell dose impact on overall survival (OS) has been investigated in various studies with contradicting conclusions. We utilized machine learning (ML) techniques to predict alloHCT outcomes based on patient and transplant-related characteristics, using single-center data to identify critical and often overlooked variable interactions. Our predictive model improves on the conventional Cox regression analysis with the ability to capture complex high-dimensional relationships. The application of SHAP (SHapley Additive exPlanations), an explainable artificial intelligence (XAI) technique, allowed us to discern new and clinically relevant feature-outcome relationships. The cohort used for developing the machine learning model covers 1153 patients across different diagnoses, transplanted between 2010 to 2019. In particular, we chose XGBoost, a tree-based ML model, that is known for its high performance with tabular data. After further investigation by SHAP, we narrowed the study cohort to 674 acute leukemia patients (AML n=549, ALL n=125) with a median age of 54 (18-76), where we identified clear interaction between CD34+ cell dose of the peripheral blood stem cell (PBSC) grafts and patient age at alloHCT for acute leukemia patients (Figure 1). Three cell dose groups (low, medium and high) were determined using maximally selected rank statistics method, separately for the younger (≤45-years-old) and older (>45-year-old) patient age groups. On univariate analysis for OS, within the younger cohort, the group with a lower CD34+ dose resulted in the highest 2-year OS: low vs. medium vs. high dose, 84.6% vs. 48.2% vs. 59.1% (overall p-value=0.034, Figure 2a). For the older cohort, the group with a low CD34+ dose resulted in the lowest 2-year OS: low vs. medium vs. high dose, 33.3% vs. 48.8% vs. 55.1% (overall p-value=0.054, Figure 2b). Multivariable analysis affirmed the interaction effect: for younger acute leukemia patients, a cell dose no higher than 4.3 × 106 CD34+/kg was associated with improved OS against the medium (hazard ratio [HR] 0.22, p=0.004) and higher (HR 0.25, p=0.009) cell dose groups. Conversely, for older acute leukemia patients, low CD34+ cell dose <3.8 × 106/kg was associated with worse OS against a high dose ³6.1 × 106/kg (HR 1.73, p= 0.009). The high dose group in older patients retained superior OS compared to low and medium dose groups combined (HR 0.73, p=0.024). Our findings suggest that CD34+ cell dose should be tailored by patient age for acute leukemia patients undergoing alloHCT, and XAI showcases excellent proficiency in revealing such interactions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".