Enhanced Risk Stratification for Children and Young Adults with B-Cell Acute Lymphoblastic Leukemia: A Children’s Oncology Group Report
Bibliographic record
Abstract
Abstract Current strategies to treat pediatric acute lymphoblastic leukemia rely on risk stratification algorithms using categorical data. We investigated whether using continuous variables assigned different weights would improve risk stratification. We developed and validated a multivariable Cox model for relapse-free survival (RFS) using information from 21199 patients. We constructed risk groups by identifying cutoffs of the COG Prognostic Index (PI COG ) that maximized discrimination of the predictive model. Patients with higher PI COG have higher predicted relapse risk. The PI COG reliably discriminates patients with low vs. high relapse risk. For those with moderate relapse risk using current COG risk classification, the PI COG identifies subgroups with varying 5-year RFS. Among current COG standard-risk average patients, PI COG identifies low and intermediate risk groups with 96% and 90% RFS, respectively. Similarly, amongst current COG high-risk patients, PI COG identifies four groups ranging from 96% to 66% RFS, providing additional discrimination for future treatment stratification. When coupled with traditional algorithms, the novel PI COG can more accurately risk stratify patients, identifying groups with better outcomes who may benefit from less intensive therapy, and those who have high relapse risk needing innovative approaches for cure.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".