MétaCan
Menu
Back to cohort
Record W4205914074 · doi:10.1093/humrep/deab280

Predictive models of pregnancy based on data from a preconception cohort study

2021· article· en· W4205914074 on OpenAlexaboutno aff
Jennifer J. Yland, Taiyao Wang, Zahra Zad, Sydney K. Willis, Tanran R. Wang, Amelia K. Wesselink, Tammy Jiang, Elizabeth E. Hatch, Lauren A. Wise, Ioannis Ch. Paschalidis

Bibliographic record

VenueHuman Reproduction · 2021
Typearticle
Languageen
FieldMedicine
TopicOvarian function and disorders
Canadian institutionsnot available
FundersPrecursory Research for Embryonic Science and TechnologyEunice Kennedy Shriver National Institute of Child Health and Human DevelopmentNational Center for Advancing Translational SciencesNational Institute of General Medical SciencesNational Institutes of HealthNational Science Foundation
KeywordsPregnancyLogistic regressionFertilityCohortInfertilityMedicineCohort studyReceiver operating characteristicDemographyObstetricsPsychologyGynecologyPopulationEnvironmental healthBiologyInternal medicine

Abstract

fetched live from OpenAlex

STUDY QUESTION: Can we derive adequate models to predict the probability of conception among couples actively trying to conceive? SUMMARY ANSWER: Leveraging data collected from female participants in a North American preconception cohort study, we developed models to predict pregnancy with performance of ∼70% in the area under the receiver operating characteristic curve (AUC). WHAT IS KNOWN ALREADY: Earlier work has focused primarily on identifying individual risk factors for infertility. Several predictive models have been developed in subfertile populations, with relatively low discrimination (AUC: 59-64%). STUDY DESIGN, SIZE, DURATION: Study participants were female, aged 21-45 years, residents of the USA or Canada, not using fertility treatment, and actively trying to conceive at enrollment (2013-2019). Participants completed a baseline questionnaire at enrollment and follow-up questionnaires every 2 months for up to 12 months or until conception. We used data from 4133 participants with no more than one menstrual cycle of pregnancy attempt at study entry. PARTICIPANTS/MATERIALS, SETTING, METHODS: On the baseline questionnaire, participants reported data on sociodemographic factors, lifestyle and behavioral factors, diet quality, medical history and selected male partner characteristics. A total of 163 predictors were considered in this study. We implemented regularized logistic regression, support vector machines, neural networks and gradient boosted decision trees to derive models predicting the probability of pregnancy: (i) within fewer than 12 menstrual cycles of pregnancy attempt time (Model I), and (ii) within 6 menstrual cycles of pregnancy attempt time (Model II). Cox models were used to predict the probability of pregnancy within each menstrual cycle for up to 12 cycles of follow-up (Model III). We assessed model performance using the AUC and the weighted-F1 score for Models I and II, and the concordance index for Model III. MAIN RESULTS AND THE ROLE OF CHANCE: Model I and II AUCs were 70% and 66%, respectively, in parsimonious models, and the concordance index for Model III was 63%. The predictors that were positively associated with pregnancy in all models were: having previously breastfed an infant and using multivitamins or folic acid supplements. The predictors that were inversely associated with pregnancy in all models were: female age, female BMI and history of infertility. Among nulligravid women with no history of infertility, the most important predictors were: female age, female BMI, male BMI, use of a fertility app, attempt time at study entry and perceived stress. LIMITATIONS, REASONS FOR CAUTION: Reliance on self-reported predictor data could have introduced misclassification, which would likely be non-differential with respect to the pregnancy outcome given the prospective design. In addition, we cannot be certain that all relevant predictor variables were considered. Finally, though we validated the models using split-sample replication techniques, we did not conduct an external validation study. WIDER IMPLICATIONS OF THE FINDINGS: Given a wide range of predictor data, machine learning algorithms can be leveraged to analyze epidemiologic data and predict the probability of conception with discrimination that exceeds earlier work. STUDY FUNDING/COMPETING INTEREST(S): The research was partially supported by the U.S. National Science Foundation (under grants DMS-1664644, CNS-1645681 and IIS-1914792) and the National Institutes for Health (under grants R01 GM135930 and UL54 TR004130). In the last 3 years, L.A.W. has received in-kind donations for primary data collection in PRESTO from FertilityFriend.com, Kindara.com, Sandstone Diagnostics and Swiss Precision Diagnostics. L.A.W. also serves as a fibroid consultant to AbbVie, Inc. The other authors declare no competing interests. TRIAL REGISTRATION NUMBER: N/A.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.045
Threshold uncertainty score0.935

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.094
GPT teacher head0.320
Teacher spread0.226 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations24
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueHuman ReproductionSame topicOvarian function and disordersFrench-language works237,207