O-237 An image-based Artificial Intelligence (AI) model trained to predict blastocyst development from oocyte images correlates with key embryonic developmental parameters in a large dataset
Bibliographic record
Abstract
Abstract Study question Does MAGENTA, an AI model built to predict blastocyst development from mature oocyte images, correlate with indicators of embryonic development? Summary answer Higher MAGENTA scores correlate with blastulation and key parameters of embryonic development progression—day-3 embryo fragmentation rate, blastocyst morphology and usage fate. What is known already The oocyte is the major contributor of cytoplasmic organelles, cell membranes, and maternal mRNA that support early embryonic development. Despite the oocyte’s vital contributions to the developmental potential of blastocysts, non-invasive and objective oocyte evaluation methods are lacking. MAGENTA is a non-invasive AI tool that assesses metaphase II(MII) oocyte images, providing a score of 0-10 with higher scores indicating a higher potential for blastocyst development. While MAGENTA was only trained to predict blastulation, its scores have shown correlation with other indicators of embryo development, highlighting the contribution of oocyte quality to the success of embryonic development and the overall cycle. Study design, size, duration This validation study included a dataset of 15,521 images of fresh denuded MII oocytes immediately post-ICSI (6964 donor, 8557 patient oocytes; age: 18-48 years from a Spanish fertility clinic. 7734 oocytes developed into blastocysts while 7787 oocytes did not (2458 non-fertilized, 611 abnormally fertilized). Oocyte images and accompanying clinical parameters were assessed by MAGENTA to produce scores (0-10), indicating chances of blastocyst development, which were also divided into four groups (0-2.5,2.6-5, 5.1-7.5, 7.6-10) for analysis. Participants/materials, setting, methods MAGENTA scores were compared to true blastocyst development outcomes and analyzed for correlation with key embryo development parameters—day-3 fragmentation percentage, blastocyst development, blastocyst quality, and blastocyst usage fate as decided by embryologists. Embryologist-assigned Gardner grading included 1079 excellent-quality (ICM and TE=A), 4940 good-quality (ICM/TE=B), 1264 fair-quality (ICM/TE=C), and 376 poor-quality (ICM/TE=D) blastocysts. Correlation analyses were assessed by Welch’s t-test, One-Way Analysis of Variance (ANOVA) with Tukey’s post-hoc pairwise comparisons, or Two Proportions z-test. Main results and the role of chance Oocyte MAGENTA scores decreased significantly in a stepwise manner with higher day-3 embryo fragmentation rates(p < 0.01), except between 11-25% and 26-35% groups(p = 0.99) (≤10%[6.5, n = 10286], 11-25%[5.5, n = 1696], 26-35%[5.6, n = 178], >35%[4.5, n = 149]). A similar relationship was found when assessing only oocytes that developed into blastocysts(≤10%[6.9, n = 6907], 11-25%[6.2, n = 716], 26-35%[6.7, n = 48], >35%[4.1, n = 22])(p < 0.01), except between 11-25% and 26-35% groups(p = 0.64). Oocytes that developed into blastocysts also had significantly higher MAGENTA scores than those that did not(6.9 vs. 5.1, p < 0.0001), with a stepwise increase in the proportion of blastocysts developed within each sequential MAGENTA score group (26%[n = 3058], 44%[n = 2855], 52%[n = 3323], 63%[n = 6285]). Subgroup analysis revealed an increase in oocyte MAGENTA scores with increasing blastocyst quality: Non-blastocysts vs. Poor (5.1 vs. 6, p < 0.0001), Poor vs. Fair (6 vs. 6.2, p = 0.62), Fair vs. Good (6.2 vs. 6.9, p < 0.0001), Good vs. Excellent (6.9 vs. 7.5, p < 0.0001). Similarly, there was a significant stepwise increase in the proportion of Excellent blastocysts (2.1%, 4%, 7%, 11%; p < 0.0001) and Good blastocysts (15%, 28%, 34%, 41%; p < 0.0001) within sequentially increasing MAGENTA score groups. Additionally, MAGENTA scores correlated with blastocyst usage decisions by embryologists, with significantly higher MAGENTA scores for oocyte that became embryos selected for transfer(7.2,n=407) or cryopreservation (6.9,n=6733) than those discarded (5.2,n=8379; both p < 0.0001). Limitations, reasons for caution Analysis of day-3 embryo fragmentation was limited by small sample sizes of 26-35% and >35% groups. Poor-quality sample size was limited compared to other quality groups. Additional data is needed to further elucidate any quality differences between oocytes that develop into embryos chosen for fresh transfers and for cryopreservation. Wider implications of the findings MAGENTA’s oocyte assessments correlate with key parameters of embryonic development, such as fragmentation, blastocyst development, quality, and utilization, which are linked to implantation potential. These results highlight the critical role of oocytes in driving embryonic development and MAGENTA’s potential to offer valuable insights into cycle potential from the oocyte stage. Trial registration number No
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".