MétaCan
Menu
← Back to cohort
Record W4411399610 · doi:10.2196/69220

Development of Machine Learning–Based Risk Prediction Models to Predict Rapid Weight Gain in Infants: Analysis of Seven Cohorts

2025· article· en· W4411399610 on OpenAlexvenueno aff
Miaobing Zheng, Yuxin Zhang, Rachel Laws, Peter Vuillermin, Jodie M Dodd, Li Ming Wen, Louise A. Baur, Rachael W. Taylor, Rebecca Byrne, Anne‐Louise Ponsonby, Kylie D. Hesketh

Bibliographic record

VenueJMIR Public Health and Surveillance · 2025
Typearticle
Languageen
FieldMedicine
TopicGestational Diabetes Research and Management
Canadian institutionsnot available
Fundersnot available
KeywordsMedicineReceiver operating characteristicBreastfeedingBirth weightFalse positive paradoxMachine learningPredictive modellingWeight gainArtificial intelligenceStatisticsPregnancyPediatricsComputer scienceMathematicsBody weightInternal medicine

Abstract

fetched live from OpenAlex

Background: Rapid weight gain (RWG) during infancy, defined as an upward crossing of one centile line on a weight growth chart, is highly predictive of subsequent obesity risk. Identification of infant RWG could facilitate obesity risk assessment from infancy. Objective: Leveraging machine learning (ML) algorithms, this study aimed to develop and validate risk prediction models to identify infant RWG by the age of 1 year. Methods: Data from 7 Australian and New Zealand cohorts were pooled for risk model development and validation (n=5233). A total of 8 ML algorithms predicted infant RWG using routinely available prenatal and early postnatal factors, including maternal prepregnancy weight status, maternal smoking during pregnancy, gestational age, parity, infant sex, birth weight, any breastfeeding and timing of solids introduction at the age of 6 months. Pooled data were randomly split into a training dataset (70%) and a test dataset (30%) for model training and validation, respectively. Model consistency was evaluated using 5-fold cross-validation. Model predictive performance was evaluated by area under the receiver operating characteristic (ROC) curve (AUC), accuracy, precision, sensitivity, specificity, and Cohen κ. Results: The average prevalence of infant RWG was 27%. In the training dataset, all ML algorithms showed acceptable to excellent discrimination with AUCs ranging from 0.75 to 0.86. Accuracy, which indicates the overall correctness of the model, ranged from 0.69 to 0.78. Precision, which measures the model's ability to avoid false positives, ranged from 0.68 to 0.77. The spread of sensitivity, specificity, and Cohen κ of all models was 0.68-0.80, 0.65-0.78, and 0.38-0.56, respectively. Of the 8 algorithms, the Gradient Boosting model showed the most favorable predictive accuracy. Validation of the Gradient Boosting model in the testing dataset exhibited excellent discrimination (AUC 0.3-0.6) and good ability to make accurate predictions, particularly true positive cases (with accuracy and sensitivity>0.75), but modest performance for precision (0.57-0.60) and Cohen κ (0.47-0.52). Conclusions: This study developed the first set of ML-based risk prediction models to identify infants' risk of experiencing RWG by the age of 1 year with acceptable accuracy. The models could be feasibly integrated into routine child growth monitoring and may facilitate population-wide early obesity risk assessment in primary health care.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.020
metaresearch head score (Gemma)0.025
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.020
Threshold uncertainty score0.106

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0200.025
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.003
Bibliometrics0.0020.001
Science and technology studies0.0010.000
Scholarly communication0.0010.001
Open science0.0010.002
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.021
GPT teacher head0.306
Teacher spread0.285 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Public Health and Surveillance→Same topicGestational Diabetes Research and Management→French-language works237,207→