Response to the Commentary ‘Causes of ART‐related outcomes in the COVID‐19 era’
Bibliographic record
Abstract
We thank Bruckner and Gemmill for their commentary1 on the methodological issues related to our paper about birth outcomes following the use of assisted reproductive technology (ART) during the COVID-19 pandemic.2 Their criticism was primarily directed at two of the statistical methods we used to assess temporal trends prior to the pandemic and changes following the pandemic onset. The first concern was that an inappropriate reference group was used to assess changes in demographic characteristics of women with singleton and multiple births in the United States in 2015–2020. Bruckner and Gemmill state that we incorrectly used a ‘stacked calendar’ approach, with the reference group used being the average across all women who conceived in the month of March in each year between 2015 and 2019. However, as was noted in the Methods section of our paper, and reiterated in the footnote to Table 2, the results were estimated via interrupted time series models (ITS). Interrupted time series models account for temporal trends in the pre-interruption period, which as Brucker and Gemmill correctly point out, is needed for a proper comparison.3 The second criticism from Bruckner and Gemmill was directed at the ‘implausibly narrow’ confidence intervals derived from our segmented Poisson model for the temporal changes in stillbirth rates among women who conceived by ART. Their point was supported by the results from an ARIMA model, which they fit to the data from our study, and which showed wider intervals. Unfortunately, their analysis confuses two related concepts, namely confidence intervals and prediction intervals. In their commentary, Bruckner and Gemmill refer to confidence intervals in the text, but what they present were prediction intervals. In general, confidence intervals represent the range for the expected value of the outcome given specific covariates, which in this context is centred around the combination of regression coefficients from the fitted model. In contrast, prediction intervals give an estimated range for a future observation (not the expected value). Regression models estimate the mean response more accurately than they predict a future value and therefore confidence intervals will always be narrower than prediction intervals. We appreciate that prediction intervals lead to a more cautious interpretation of our results, although both approaches confirm the primary findings of our study that show a spike in stillbirth rates in December 2020 among women who conceived following ART. The mechanics of the observed increase in stillbirths in our study is unclear, however, and Bruckner and Gemmill argue that stillbirths among women who conceived in March would have occurred before December 2020 because a majority of stillbirths occur prior to 28 weeks' gestation. However, it should be noted that only 52%–53% of stillbirths occurred at 20–27 weeks' gestation in 2015–2020 and in prior years.4 The calendar month impacted by conceptions in March 2020 depends on the gestational age at which foetal deaths increased, any lag period between gestational age at foetal death (in utero) and stillbirth (a birth following foetal death), and the gestational age distribution of live births conceived in March. We examined this issue further by performing additional analysis using an ITS model to compare conception cohorts among women who conceived following ART (Figure 1). The results show that the estimated conceptions in March 2020 were associated with an increased rate of stillbirth (rate ratio 3.38, 95% confidence interval 1.39, 7.02) compared with expected rates derived from the modelled temporal trend between 2015 and 2019. While both birth and conception cohort approaches yielded similar results, our explanations for the increase in stillbirths in December 2020 remain speculative. Lack of data on miscarriages possibly resulting in left-truncation bias, and other limitations of the data set are other possible explanations. We appreciate the various approaches to analyses and the ongoing dialogue in the quest for improving statistical methods and their application to gain deeper insight into temporal changes in maternal characteristics and perinatal outcomes. This research was supported by funding from the Canadian Institutes of Health Research (grant number F17-02161). MAB received Grants from Ferring Pharmaceutical and Abbvie Pharmaceuticals. None of the other authors have financial relationships relevant to this article to disclose. None. The authors report no conflict of interest. SL and JNB drafted the letter to the editor. JNB conducted data analysis. GMM, NR, AB, JSB, MAB, CVA, and KSJ contributed to the intelectual content and the final version of the letter to the editor. All data used in this study are publicly available in deidentified form at https://www.cdc.gov/nchs/data_access/vitalstatsonline.htm
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.019 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".