Prognostic Accuracy of BPD Definitions for Long-Term Outcomes in Preterm Infants: A Systematic Review
Bibliographic record
Abstract
BACKGROUND AND OBJECTIVES: Since the first description of bronchopulmonary dysplasia (BPD), multiple definitions to diagnose BPD and its grading have been published. Several studies have compared the predictive performance of these definitions for long-term outcomes. The objective was to identify the BPD definition with the optimal predictive performance for long-term respiratory and neurological outcomes in preterm infants. METHODS: An electronic search identified studies in Medline and Embase from inception to August 2024. Studies assessing the performance of one or more BPD definitions for predicting long-term respiratory and/or neurological outcomes were included. We used the Quality in Prognostic Studies (QUIPS) tool for bias assessment. Reported prognostic accuracy of 5 BPD definitions (the 1988 Shennan, the 2001 National Institutes of Health [NIH], the 2017 Canadian Neonatal Network, the 2018 NIH, and the 2019 Neonatal Research Network definition) was tabulated using specificity, sensitivity, C statistic, risk, or odds ratio. RESULTS: Of the 6045 identified studies, 18 were included. Heterogeneity between studies resulted in inconsistent prognostic accuracy for long-term outcomes. The 2001 NIH definition showed higher prognostic accuracy for respiratory and neurological outcomes compared with the 1988 Shennan BPD definition. Only 5 studies showed a low to moderate risk of bias, and a sensitivity analysis confirmed the results. The limitations included challenges in comparing studies due to population heterogeneity and outcome definitions. CONCLUSIONS: This systematic review shows that comparisons between the 2001 NIH definition and newer BPD definitions yield inconsistent results for predicting long-term outcomes. None of the current BPD definitions consistently provided sufficient prognostic accuracy for long-term respiratory and neurodevelopmental sequelae in very preterm infants.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.108 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.008 | 0.010 |
| Bibliometrics | 0.011 | 0.010 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".