MétaCan
Menu
Retour à la cohorte
Enregistrement W4386023761 · doi:10.1111/1471-0528.17641

Performance of <scp>ChatGPT</scp> in medical examinations: A systematic review and a meta‐analysis

2023· review· en· W4386023761 sur OpenAlexaff
Gabriel Levin, Nir Horesh, Yoav Brezinov, Raanan Meyer

Notice bibliographique

RevueBJOG An International Journal of Obstetrics & Gynaecology · 2023
Typereview
Langueen
DomaineMedicine
ThématiqueArtificial Intelligence in Healthcare and Education
Établissements canadiensMcGill UniversityJewish General Hospital
Organismes subventionnairesnon disponible
Mots-clésScopusWeb of scienceEnglish languageMeta-analysisConversationSystematic reviewMEDLINEMultiple choiceComputer scienceMedical educationPsychologyMedicineMathematics educationPathologyInternal medicine

Résumé

récupéré en direct d'OpenAlex

The use of ChatGPT, an artificial intelligence (AI) language model, has been described in various scientific and medical applications.1 With its human-like conversation capacity and large quantity of data, ChatGPT has the potential to become an important medical education tool.2 ChatGPT's performance in different medical knowledge examinations has been recently studied in various medical disciplines. However, reported rates of correct answers vary dramatically across different examinations and medical fields.3, 4 We aimed to conduct a meta-analysis of studies reporting ChatGPT's performance in medical examinations with multiple-choice questions. PubMed, Scopus and Web of Science websites were searched for relevant articles from the inception of these databases to 2 June 2023, by use of the word “ChatGPT” (no equivalent Mesh term exists). We manually reviewed every article title and abstract. In case an abstract was not available, we accessed the abstract on the journal's web site. We included all peer-reviewed articles, assessing the performance (number of right answers/number of questions) of ChatGPT in multiple-choice questions in the field of medicine. Exclusion criteria were: (i) performance of ChatGPT not in English language; (ii) a study evaluating performance of ChatGPT in a setting other than multiple-choice question (e.g. open questions, frequently asked questions); and (iii) a study not reporting ChatGPT version 3.5. All review stages were conducted independently by two reviewers (RM and GL). Disagreements were resolved by discussion with a third reviewer (YB). Data were extracted from each included study without modifications and a database was constructed containing the study field of medicine (e.g. dermatology, plastic surgery etc.), cohort size (number of questions answered), number of correct questions, ChatGPT performance (number of correct answers/numbers of answered questions) and 95% CI for performance rate. We used MedCalc Statistical Software version 19.2.6 (MedCalc Software bv, Ostend, Belgium) and OpenMeta[Analyst] for the analysis. The process of literature search and article selection is presented in the Supplementary material. Finally, a total of 19 articles were included in the analysis. Two articles (11%) studied Plastic Surgery examinations, two (11%) studied the United Stated Medical Licensing Examinations (USMLE) and two (11%) studied anaesthesia examinations. All other publications studied different medical fields (Table 1). The median number of questions per examination was 242, ranging from 20 in medical physiology examinations to 3705 in anaesthesia examinations (mean 524.5 ± 847.2 standard deviation, Figure 1, Table 1). Overall performance of ChatGPT ranged from 40% in the biomedical admission test to 100% in a diabetes knowledge questionnaire. The mean performance of ChatGPT was 61.1% (95% CI 56.1%–66.0%. As the literature regarding the performance of ChatGPT in medical education is mounting, summarising the current performance of ChatGPT in medical examinations can provide an insight into the present and future of AI in medical education. Our meta-analysis suggests that ChatGPT correctly answered the majority of multiple-choice questions in medical examinations and demonstrated a performance of approximately a passing grade. Medical education, test preparation services, and medical examinations form a large industrial market.5 Currently, it seems that the use of ChatGPT for examination preparation should be prudent, and preparations for examinations using this platform should be done with the proper caution. A correct response rate of more than 95% may allow ChatGPT to become a reliable educational tool,6 and it is unknown whether future versions will reach this maturity level, as the training data set is not specifically developed with a focus on medical education. Our limitations are the inclusion of only ChatGPT version 3.5 studies, the heterogeneity of the included studies with unreported numbers of choices per examination, and the possible overestimation of the tool performance. Future meta-analyses may include future versions of AI chatbots to provide updated understanding of their role in medical education. None. This research received no external funding. None declared. No ethical approval was needed by the institutional review board as this study analyzed only published available public data and no human patients data was sused. The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions. Data S1. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,003
score de la tête « metaresearch » (Gemma)0,029
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Méta-analyse · Signal consensuel: aucune
GenreSignal candidat: Synthèse · Signal consensuel: Synthèse
Score de désaccord entre enseignants0,839
Score d'incertitude au seuil0,979

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0030,029
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0050,001
Bibliométrie0,0040,002
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0010,000
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,238
Tête enseignante GPT0,477
Écart entre enseignants0,239 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeMéta-analyse
Domainenon disponible
GenreSynthèse

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations100
Publié2023
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueBJOG An International Journal of Obstetrics & GynaecologyMême sujetArtificial Intelligence in Healthcare and EducationTravaux en français237 207