Benefits and challenges with diagnosing chronic and late acute GVHD in children using the NIH consensus criteria
Bibliographic record
Abstract
Chronic graft-versus-host disease (cGVHD) and late acute graft-versus-host disease (L-aGVHD) are understudied complications of allogeneic hematopoietic stem cell transplantation in children. The National Institutes of Health Consensus Criteria (NIH-CC) were designed to improve the diagnostic accuracy of cGVHD and to better classify graft-versus-host disease (GVHD) syndromes but have not been validated in patients <18 years of age. The objectives of this prospective multi-institution study were to determine: (1) whether the NIH-CC could be used to diagnose pediatric cGVHD and whether the criteria operationalize well in a multi-institution study; (2) the frequency of cGVHD and L-aGVHD in children using the NIH-CC; and (3) the clinical features and risk factors for cGVHD and L-aGVHD using the NIH-CC. Twenty-seven transplant centers enrolled 302 patients <18 years of age before conditioning and prospectively followed them for 1 year posttransplant for development of cGVHD. Centers justified their cGVHD diagnosis according to the NIH-CC using central review and a study adjudication committee. A total of 28.2% of reported cGVHD cases was reclassified, usually as L-aGVHD, following study committee review. Similar incidence of cGVHD and L-aGVHD was found (21% and 24.7%, respectively). The most common organs involved with diagnostic or distinctive manifestations of cGVHD in children include the mouth, skin, eyes, and lungs. Importantly, the 2014 NIH-CC for bronchiolitis obliterans syndrome perform poorly in children. Past acute GVHD and peripheral blood grafts are major risk factors for cGVHD and L-aGVHD, with recipients ≥12 years of age being at risk for cGVHD. Applying the NIH-CC in pediatrics is feasible and reliable; however, further refinement of the criteria specifically for children is needed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".