External validation of prognostic models to predict stillbirth using International Prediction of Pregnancy Complications (<scp>IPPIC</scp>) Network database: individual participant data meta‐analysis
Bibliographic record
Abstract
OBJECTIVE: Stillbirth is a potentially preventable complication of pregnancy. Identifying women at high risk of stillbirth can guide decisions on the need for closer surveillance and timing of delivery in order to prevent fetal death. Prognostic models have been developed to predict the risk of stillbirth, but none has yet been validated externally. In this study, we externally validated published prediction models for stillbirth using individual participant data (IPD) meta-analysis to assess their predictive performance. METHODS: MEDLINE, EMBASE, DH-DATA and AMED databases were searched from inception to December 2020 to identify studies reporting stillbirth prediction models. Studies that developed or updated prediction models for stillbirth for use at any time during pregnancy were included. IPD from cohorts within the International Prediction of Pregnancy Complications (IPPIC) Network were used to validate externally the identified prediction models whose individual variables were available in the IPD. The risk of bias of the models and cohorts was assessed using the Prediction study Risk Of Bias ASsessment Tool (PROBAST). The discriminative performance of the models was evaluated using the C-statistic, and calibration was assessed using calibration plots, calibration slope and calibration-in-the-large. Performance measures were estimated separately in each cohort, as well as summarized across cohorts using random-effects meta-analysis. Clinical utility was assessed using net benefit. RESULTS: Seventeen studies reporting the development of 40 prognostic models for stillbirth were identified. None of the models had been previously validated externally, and the full model equation was reported for only one-fifth (20%, 8/40) of the models. External validation was possible for three of these models, using IPD from 19 cohorts (491 201 pregnant women) within the IPPIC Network database. Based on evaluation of the model development studies, all three models had an overall high risk of bias, according to PROBAST. In the IPD meta-analysis, the models had summary C-statistics ranging from 0.53 to 0.65 and summary calibration slopes ranging from 0.40 to 0.88, with risk predictions that were generally too extreme compared with the observed risks. The models had little to no clinical utility, as assessed by net benefit. However, there remained uncertainty in the performance of some models due to small available sample sizes. CONCLUSIONS: The three validated stillbirth prediction models showed generally poor and uncertain predictive performance in new data, with limited evidence to support their clinical application. The findings suggest methodological shortcomings in their development, including overfitting. Further research is needed to further validate these and other models, identify stronger prognostic factors and develop more robust prediction models. © 2021 The Authors. Ultrasound in Obstetrics & Gynecology published by John Wiley & Sons Ltd on behalf of International Society of Ultrasound in Obstetrics and Gynecology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".