P-735 Is Artificial Intelligence (AI) currently able to provide evidence-based scientific responses on methods that can improve the outcomes of embryo transfers?
Notice bibliographique
Résumé
Abstract Study question Could chatbots potentially be used as tools for both patients’ and physicians’ education in an infertility journey? Summary answer Artificial Intelligence (AI) is not yet in a position to give clear, evidence-based recommendations in the field of fertility, particularly concerning embryo transfer. What is known already The rapid development of AI has raised questions about its potential uses in different sectors of everyday life. Individuals from diverse fields have tried to incorporate its usage in their personal but also in their professional lives, with varying levels of success. Specifically in medicine, ChatGPT has been studied intensely with more than 400 related articles having been indexed in PubMed until May 2023, while AI has reportedly also succeeded in several standardized medical tests and passed several Medical Board examinations. Hence the question arises whether chatbots could be used as tools for clinical decision-making or patients’ and physicians’ education. Study design, size, duration We used nine of the most popular free AI chatbots available (ChatGPT, Bard, Writesonic, You, Perplexity, Learnt, Bing, Magickpen, and Rytr) and entered the following command: “Write me a 300-word scientific essay about evidence-based methods that can improve the outcomes of embryo transfer” in May 2023. We collected the responses and extracted the methods each chatbot suggested. When sufficient similarity among answers was present we categorized the answers under one category to facilitate the study. Participants/materials, setting, methods We calculated descriptive statistics and the prevalence of each response. We used as a comparator for widely acceptable practices that are proven to improve embryo transfer outcomes the 2017 ASRM guideline on performing embryo transfer taking also into consideration more current literature. Data was compared using chi-squared tests and a level of p < 0.05 was used as the threshold of statistical significance. Main results and the role of chance The range of the recommendations was from one to nine per chatbot (median = 4, IQR = 2) with an average of 4.78 suggestions per chatbot. Out of a total of 43 recommendations, which could be grouped into 19 similar categories, only 3/19 (15.8%) were evidence-based practices, those being “ultrasound-guided embryo transfer”, which was also the most commonly appearing response, appearing in 7/9 (77.8%) chatbots, “single embryo transfer” appearing in 4/9 (44.4%) chatbots and “use of a soft catheter” in 2/9 (22.2%) chatbots. Some controversial responses appeared even more often than the two latter evidence-based ones with “preimplantation genetic testing (PGT)” being the second most common response with 6/9 chatbots suggesting it (66.7%) and the vague suggestion of “optimal endometrium preparation” being in the third place with 5/9 chatbots (55.6%), both non-evidence-based practices to improve embryo transfer. The majority of the recommendations were unique, with 10 answers appearing only once (1/9; 11.1%), those being “endometrial scratching”, “natural cycle embryo transfer”, “use of GnRH antagonist protocol”, “blastocyst transfer”, “use of specialized catheter”, “preimplantation genetic screening (PGS)”, “mock embryo transfer”, “uninterrupted embryo culture”, “minimizing transfer time”, and “maintaining temperature and pH of the culture media”. Limitations, reasons for caution Firstly, we only used some of the most popular chatbots and not all existing ones. Additionally, chatbots tend to give different answers when repeatedly asked the same questions. However, we believe that our study effectively captures the prevailing practices of chatbot users, who typically pose a question only once. Wider implications of the findings Both patients and physicians should be wary of guiding care based on chatbot recommendations in infertility since the majority of responses consists of scientifically unsupported recommendations. Chatbot results might improve with time especially if trained to obtain information from validated medical databases, however, this will have to be verified scientifically. Trial registration number N/A
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».