MétaCan
Menu
Back to cohort
Record W4400299350 · doi:10.1093/humrep/deae108.1060

P-735 Is Artificial Intelligence (AI) currently able to provide evidence-based scientific responses on methods that can improve the outcomes of embryo transfers?

2024· article· en· W4400299350 on OpenAlexaff
A. Kolokythas, Marc Dahan

Bibliographic record

VenueHuman Reproduction · 2024
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsMcGill University
Fundersnot available
KeywordsEmbryoArtificial intelligenceComputer scienceBiologyComputational biologyGenetics

Abstract

fetched live from OpenAlex

Abstract Study question Could chatbots potentially be used as tools for both patients’ and physicians’ education in an infertility journey? Summary answer Artificial Intelligence (AI) is not yet in a position to give clear, evidence-based recommendations in the field of fertility, particularly concerning embryo transfer. What is known already The rapid development of AI has raised questions about its potential uses in different sectors of everyday life. Individuals from diverse fields have tried to incorporate its usage in their personal but also in their professional lives, with varying levels of success. Specifically in medicine, ChatGPT has been studied intensely with more than 400 related articles having been indexed in PubMed until May 2023, while AI has reportedly also succeeded in several standardized medical tests and passed several Medical Board examinations. Hence the question arises whether chatbots could be used as tools for clinical decision-making or patients’ and physicians’ education. Study design, size, duration We used nine of the most popular free AI chatbots available (ChatGPT, Bard, Writesonic, You, Perplexity, Learnt, Bing, Magickpen, and Rytr) and entered the following command: “Write me a 300-word scientific essay about evidence-based methods that can improve the outcomes of embryo transfer” in May 2023. We collected the responses and extracted the methods each chatbot suggested. When sufficient similarity among answers was present we categorized the answers under one category to facilitate the study. Participants/materials, setting, methods We calculated descriptive statistics and the prevalence of each response. We used as a comparator for widely acceptable practices that are proven to improve embryo transfer outcomes the 2017 ASRM guideline on performing embryo transfer taking also into consideration more current literature. Data was compared using chi-squared tests and a level of p < 0.05 was used as the threshold of statistical significance. Main results and the role of chance The range of the recommendations was from one to nine per chatbot (median = 4, IQR = 2) with an average of 4.78 suggestions per chatbot. Out of a total of 43 recommendations, which could be grouped into 19 similar categories, only 3/19 (15.8%) were evidence-based practices, those being “ultrasound-guided embryo transfer”, which was also the most commonly appearing response, appearing in 7/9 (77.8%) chatbots, “single embryo transfer” appearing in 4/9 (44.4%) chatbots and “use of a soft catheter” in 2/9 (22.2%) chatbots. Some controversial responses appeared even more often than the two latter evidence-based ones with “preimplantation genetic testing (PGT)” being the second most common response with 6/9 chatbots suggesting it (66.7%) and the vague suggestion of “optimal endometrium preparation” being in the third place with 5/9 chatbots (55.6%), both non-evidence-based practices to improve embryo transfer. The majority of the recommendations were unique, with 10 answers appearing only once (1/9; 11.1%), those being “endometrial scratching”, “natural cycle embryo transfer”, “use of GnRH antagonist protocol”, “blastocyst transfer”, “use of specialized catheter”, “preimplantation genetic screening (PGS)”, “mock embryo transfer”, “uninterrupted embryo culture”, “minimizing transfer time”, and “maintaining temperature and pH of the culture media”. Limitations, reasons for caution Firstly, we only used some of the most popular chatbots and not all existing ones. Additionally, chatbots tend to give different answers when repeatedly asked the same questions. However, we believe that our study effectively captures the prevailing practices of chatbot users, who typically pose a question only once. Wider implications of the findings Both patients and physicians should be wary of guiding care based on chatbot recommendations in infertility since the majority of responses consists of scientifically unsupported recommendations. Chatbot results might improve with time especially if trained to obtain information from validated medical databases, however, this will have to be verified scientifically. Trial registration number N/A

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.816
Threshold uncertainty score0.648

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0040.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.449
GPT teacher head0.534
Teacher spread0.086 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueHuman ReproductionSame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207