An AI method to predict pregnancy loss by extracting biological indicators from embryo ultrasound recordings in early pregnancy
Bibliographic record
Abstract
B-ultrasound results are widely used in early pregnancy loss (EPL) prediction, but there are inevitable intra-observer and inter-observer errors in B-ultrasound results especially in early pregnancy, which lead to inconsistent assessment of embryonic status, and thus affect the judgment of EPL. To address this, we need a rapid and accurate model to predict pregnancy loss in the first trimester. This study aimed to construct an artificial intelligence model to automatically extract biometric parameters from ultrasound videos of early embryos and predict pregnancy loss. This can effectively eliminate the measurement error of B-ultrasound results, accurately predict EPL, and provide decision support for doctors with relatively little clinical experience. A total of 630 ultrasound videos from women with early singleton pregnancies of gestational age between 6 and 10 weeks were used for training. A two-stage artificial intelligence model was established. First, some biometric parameters such as gestational sac areas (GSA), yolk sac diameter (YSD), crown rump length (CRL) and fetal heart rate (FHR), were extract from ultrasound videos by a deep neural network named A3F-net, which is a modified neural network based on U-Net designed by ourselves. Then an ensemble learning model predicted pregnancy loss risk based on these features. Dice, IOU and Precision were used to evaluate the measurement results, and sensitivity, AUC etc. were used to evaluate the predict results. The fetal heart rate was compared with those measured by doctors, and the accuracy of results was compared with other AI models. In the biometric features measurement stage, the precision of GSA, YSD and CRL of A3F-net were 98.64%, 96.94% and 92.83%, it was the highest compared to other 2 models. Bland-Altman analysis did not show systematic deviations between doctors and AI. The mean and standard deviation of the mean relative error between doctors and the AI model was 0.060 ± 0.057. In the EPL prediction stage, the ensemble learning models demonstrated excellent performance, with CatBoost being the best-performing model, achieving a precision of 98.0% and an AUC of 0.969 (95% CI: 0.962-0.975). In this study, a hybrid AI model to predict EPL was established. First, a deep neural network automatically measured the biometric parameters from ultrasound video to ensure the consistency and accuracy of the measurements, then a machine learning model predicted EPL risk to support doctors making decisions. The use of our established AI model in EPL prediction has the potential to assist physicians in making more accurate and timely clinical decision in clinical application.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".