External validation of the RISC, RISC-Malawi, and PERCH clinical prediction rules to identify risk of death in children hospitalized with pneumonia
Bibliographic record
Abstract
BACKGROUND: Existing scores to identify children at risk of hospitalized pneumonia-related mortality lack broad external validation. Our objective was to externally validate three such risk scores. METHODS: We applied the Respiratory Index of Severity in Children (RISC) for HIV-negative children, the RISC-Malawi, and the Pneumonia Etiology Research for Child Health (PERCH) scores to hospitalized children in the Pneumonia REsearch Partnerships to Assess WHO REcommendations (PREPARE) data set. The PREPARE data set includes pooled data from 41 studies on pediatric pneumonia from across the world. We calculated test characteristics and the area under the curve (AUC) for each of these clinical prediction rules. RESULTS: The RISC score for HIV-negative children was applied to 3574 children 0-24 months and demonstrated poor discriminatory ability (AUC = 0.66, 95% confidence interval (CI) = 0.58-0.73) in the identification of children at risk of hospitalized pneumonia-related mortality. The RISC-Malawi score had fair discriminatory value (AUC = 0.75, 95% CI = 0.74-0.77) among 17 864 children 2-59 months. The PERCH score was applied to 732 children 1-59 months and also demonstrated poor discriminatory value (AUC = 0.55, 95% CI = 0.37-0.73). CONCLUSIONS: In a large external application of the RISC, RISC-Malawi, and PERCH scores, a substantial number of children were misclassified for their risk of hospitalized pneumonia-related mortality. Although pneumonia risk scores have performed well among the cohorts in which they were derived, their performance diminished when externally applied. A generalizable risk assessment tool with higher sensitivity and specificity to identify children at risk of hospitalized pneumonia-related mortality may be needed. Such a generalizable risk assessment tool would need context-specific validation prior to implementation in that setting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".