Application of CIBMTR risk score to NIH chronic GVHD at individual centers
Bibliographic record
Abstract
To the editor: A new risk score to predict mortality in patients with chronic graft-versus-host disease (GVHD) was recently reported by Arora et al by analyzing a large amount of data between 1995 and 2004 from the Center for International Blood and Marrow Transplant Registry (CIBMTR).1 The risk score consists of 10 variables defined at transplantation or at onset of chronic GVHD that are objective and easy to evaluate at any transplant center. Because performance of this risk score has not been examined in contemporary patients with chronic GVHD defined by the National Institutes of Health (NIH) criteria2,3 and because experience with application of the score to patients at individual centers is limited, we examined performance of the risk score in 376 consecutive patients with leukemia or myelodysplastic syndrome who received initial systemic treatment of NIH chronic GVHD between 2006 and 2010 at 2 individual centers (Figure 1). Patients had given written consent allowing the use of medical records for research in accordance with the Declaration of Helsinki, and the institutional review boards at the Fred Hutchinson Cancer Research Center (FHCRC) and the Princess Margaret Hospital (PMH) approved the study. Figure 1 Application of the CIBMTR risk score to patients with systemically treated NIH chronic GVHD between 2006 and 2010 at 2 individual centers. Data represented includes (A) proportions of risk groups, (B) overall survival, and (C) non-relapse mortality. Diseases ... Compared with CIBMTR’s study, this contemporary cohort included higher proportions of patients ≥60 years of age and patients who received tacrolimus or T-cell depletion as GVHD prophylaxis and decreased proportions of patients with <5 months from transplantation to onset of chronic GVHD, hyperbilirubinemia, or thrombocytopenia. In contrast to the CIBMTR’s study, the current cohorts contained few patients in risk group 4 and no patients in risk group 5 or 6. In addition, most patients (70-80%) were classified as risk group 2 (Figure 1A). Overall survival (OS) was well stratified according to risk groups at both centers (Figure 1B), whereas nonrelapse mortality (NRM) was not well stratified at the FHCRC (Figure 1C). OS for risk group 1 was similar to CIBMTR results at both centers. Compared with CIBMTR results, OS for risk group 2 was slightly higher at FHCRC and slightly lower at PMH, and OS for risk group 3 was higher at FHCRC and much lower at PMH, despite favorable demographics at PMH. Compared with CIBMTR results, NRM for risk group 1 was slightly higher at FHCRC and similar at PMH, NRM for risk group 2 was similar at both centers, and NRM for risk group 3 was lower at FHCRC and much higher at PMH. Our results confirm that the CIBMTR risk score performs well in predicting differences in OS in contemporary patients treated for NIH chronic GVHD. On the other hand, we identified at least 3 caveats in applying the risk score to our patients: (1) the proportion of patients with risk groups 4 to 6 was very low, (2) a finer separation of risk group 2 might be helpful, and (3) factors accounting for the center-specific difference in mortality, particularly for patients in risk group 3, remain to be determined. The dedicated long-term follow-up program at FHCRC could have contributed to the better survival of patients in this category.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.028 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".