MétaCan
Menu
Back to cohort
Record W4409696667 · doi:10.1093/humrep/deaf066

An interpretable artificial intelligence approach to differentiate between blastocysts with similar or same morphological grades

2025· article· en· W4409696667 on OpenAlexaffabout
Hang Liu, L.M. Chen, Guanqiao Shan, Chen Sun, Changfu Lu, Hongqing Liao, Shuoping Zhang, X. X. Xu, Qiuyun Yan, Fei Gong, Zhuoran Zhang, Changsheng Dai, Wen‐Yuan Chen, Haocong Song, Lei Chen, Shanshan Wang, Haixiang Sun, Ge Lin, Yu Sun, Yifan Gu

Bibliographic record

VenueHuman Reproduction · 2025
Typearticle
Languageen
FieldMedicine
TopicReproductive Biology and Fertility
Canadian institutionsToronto Rehabilitation InstituteUniversity of Toronto
FundersNatural Science Foundation of Hainan Province
KeywordsArtificial intelligenceBiologyAndrologyComputer scienceMachine learningMedicine

Abstract

fetched live from OpenAlex

STUDY QUESTION: Can a quantitative method be developed to differentiate between blastocysts with similar or same inner cell mass (ICM) and trophectoderm (TE) grades, while also reflecting their potential for live birth? SUMMARY ANSWER: We developed BlastScoringNet, an interpretable deep-learning model that quantifies blastocyst ICM and TE morphology with continuous scores, enabling finer differentiation between blastocysts with similar or same grades, with higher scores significantly correlating with higher live birth rates. WHAT IS KNOWN ALREADY: While the Gardner grading system is widely used by embryologists worldwide, blastocysts having similar or same ICM and TE grades cause challenges for embryologists in decision-making. Furthermore, human assessment is subjective and inconsistent in predicting which blastocysts have higher potential to result in live birth. STUDY DESIGN, SIZE, DURATION: The study design consists of three main steps. First, BlastScoringNet was developed using a grading dataset of 2760 blastocysts with majority-voted Gardner grades. Second, the model was applied to a live birth dataset of 15 228 blastocysts with known live birth outcomes to generate blastocyst scores. Finally, the correlation between these scores and live birth outcomes was assessed. The blastocysts were collected from patients who underwent IVF treatments between 2016 and 2018. For external application study, an additional grading dataset of 1455 blastocysts and a live birth dataset of 476 blastocysts were collected from patients who underwent IVF treatments between 2021 and 2023 at an external IVF institution. PARTICIPANTS/MATERIALS, SETTING, METHODS: In this retrospective study, we developed BlastScoringNet, an interpretable deep-learning model which outputs expansion degree grade and continuous scores quantifying a blastocyst's ICM morphology and TE morphology, based on the Gardner grading system. The continuous ICM and TE scores were calculated by weighting each base grade's predicted probability and summing the predicted probabilities. To represent each blastocyst's overall potential for live birth, we combined the ICM and TE scores using their odds ratios (ORs) for live birth. We further assessed the correlation between live birth rates and the ICM score, TE score, and the OR-combined score (adjusted for expansion degree) by applying BlastScoringNet to blastocysts with known live birth outcomes. To test its generalizability, we also applied BlastScoringNet to an external IVF institution, accounting for variations in imaging conditions, live birth rates, and embryologists' experience levels. MAIN RESULTS AND THE ROLE OF CHANCE: BlastScoringNet was developed using data from 2760 blastocysts with majority-voted grades for expansion degree, ICM, and TE. The model achieved mean area under the receiver operating characteristic curve values of 0.997 (SD 0.004) for expansion degree, 0.903 (SD 0.031) for ICM, and 0.943 (SD 0.040) for TE, based on predicted probabilities for each base grade. From these predicted probabilities, BlastScoringNet generated continuous ICM and TE scores, as well as expansion degree grades, for an additional 15 228 blastocysts with known live birth outcomes. Higher ICM and TE scores, along with their OR-combined scores, were significantly correlated with increased live birth rates (P < 0.0001). By fine-tuning, BlastScoringNet was applied to an external IVF institution, where higher OR-combined ICM and TE scores also significantly correlated with increased live birth rates (P = 0.00078), demonstrating consistent results across both institutions. LIMITATIONS, REASONS FOR CAUTION: This study is limited by its retrospective nature. Further prospective randomized trials are required to confirm the clinical impact of BlastScoringNet in assisting embryologists in blastocyst selection. WIDER IMPLICATIONS OF THE FINDINGS: BlastScoringNet provides an interpretable and quantitative method for evaluating blastocysts, aligned with the widely used Gardner grading system. Higher OR-combined ICM and TE scores, representing each blastocyst's overall potential for live birth, were significantly correlated with increased live birth rates. The model's demonstrated generalizability across two institutions further supports its clinical utility. These findings suggest that BlastScoringNet is a valuable tool for assisting embryologists in selecting blastocysts with the highest potential for live birth. The code and pre-trained models are publicly available to facilitate further research and widespread implementation. STUDY FUNDING/COMPETING INTEREST(S): This work was supported by the Vector Institute and the Temerty Faculty of Medicine at the University of Toronto, Toronto, Ontario, Canada, via a Clinical AI Integration Grant, and the Natural Science Foundation of Hunan Province of China (2023JJ30714). The authors declare no competing interests. TRIAL REGISTRATION NUMBER: N/A.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.576
Threshold uncertainty score0.621

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.088
GPT teacher head0.354
Teacher spread0.266 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations9
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueHuman ReproductionSame topicReproductive Biology and FertilityFrench-language works237,207