Pulmonary microbiome and transcriptome signatures reveal distinct pathobiologic states associated with mortality in two cohorts of pediatric stem cell transplant patients
Bibliographic record
Abstract
Lung injury is a major determinant of survival after pediatric hematopoietic cell transplantation (HCT). A deeper understanding of the relationship between pulmonary microbes, immunity, and the lung epithelium is needed to improve outcomes. In this multicenter study, we collected 278 bronchoalveolar lavage (BAL) samples from 229 patients treated at 32 children's hospitals between 2014-2022. Using paired metatranscriptomes and human gene expression data, we identified 4 patient clusters with varying BAL composition. Among those requiring respiratory support prior to sampling, in-hospital mortality varied from 22-60% depending on the cluster (p=0.007). The most common patient subtype, Cluster 1, showed a moderate quantity and high diversity of commensal microbes with robust metabolic activity, low rates of infection, gene expression indicating alveolar macrophage predominance, and low mortality. The second most common cluster showed a very high burden of airway microbes, gene expression enriched for neutrophil signaling, frequent bacterial infections, and moderate mortality. Cluster 3 showed significant depletion of commensal microbes, a loss of biodiversity, gene expression indicative of fibroproliferative pathways, increased viral and fungal pathogens, and high mortality. Finally, Cluster 4 showed profound microbiome depletion with enrichment of Staphylococci and viruses, gene expression driven by lymphocyte activation and cellular injury, and the highest mortality. BAL clusters were modeled with a random forest classifier and reproduced in a geographically distinct validation cohort of 57 patients from The Netherlands, recapitulating similar cluster-based mortality differences (p=0.022). Degree of antibiotic exposure was strongly associated with depletion of BAL microbes and enrichment of fungi. Potential pathogens were parsed from all detected microbes by analyzing each BAL microbe relative to the overall microbiome composition, which yielded increased sensitivity for numerous previously occult pathogens. These findings support personalized interpretation of the pulmonary microenvironment in pediatric HCT, which may facilitate biology-targeted interventions to improve outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".