MétaCan
Menu
Retour à la cohorte
Enregistrement W4378528433 · doi:10.2196/47305

Performance of the Large Language Model ChatGPT on the National Nurse Examinations in Japan: Evaluation Study

2023· article· en· W4378528433 sur OpenAlexvenueno aff
Kazuya Taira, Takahiro Itaya, Ayame Hanada

Notice bibliographique

RevueJMIR Nursing · 2023
Typearticle
Langueen
DomaineMedicine
ThématiqueArtificial Intelligence in Healthcare and Education
Établissements canadiensnon disponible
Organismes subventionnairesJapan Society for the Promotion of Science
Mots-clésMultiple choiceTest (biology)CertificationSpecialtyReading (process)Simple (philosophy)MedicineNursingPsychologyMedical educationFamily medicineLinguisticsPolitical science

Résumé

récupéré en direct d'OpenAlex

BACKGROUND: ChatGPT, a large language model, has shown good performance on physician certification examinations and medical consultations. However, its performance has not been examined in languages other than English or on nursing examinations. OBJECTIVE: We aimed to evaluate the performance of ChatGPT on the Japanese National Nurse Examinations. METHODS: We evaluated the percentages of correct answers provided by ChatGPT (GPT-3.5) for all questions on the Japanese National Nurse Examinations from 2019 to 2023, excluding inappropriate questions and those containing images. Inappropriate questions were pointed out by a third-party organization and announced by the government to be excluded from scoring. Specifically, these include "questions with inappropriate question difficulty" and "questions with errors in the questions or choices." These examinations consist of 240 questions each year, divided into basic knowledge questions that test the basic issues of particular importance to nurses and general questions that test a wide range of specialized knowledge. Furthermore, the questions had 2 types of formats: simple-choice and situation-setup questions. Simple-choice questions are primarily knowledge-based and multiple-choice, whereas situation-setup questions entail the candidate reading a patient's and family situation's description, and selecting the nurse's action or patient's response. Hence, the questions were standardized using 2 types of prompts before requesting answers from ChatGPT. Chi-square tests were conducted to compare the percentage of correct answers for each year's examination format and specialty area related to the question. In addition, a Cochran-Armitage trend test was performed with the percentage of correct answers from 2019 to 2023. RESULTS: The 5-year average percentage of correct answers for ChatGPT was 75.1% (SD 3%) for basic knowledge questions and 64.5% (SD 5%) for general questions. The highest percentage of correct answers on the 2019 examination was 80% for basic knowledge questions and 71.2% for general questions. ChatGPT met the passing criteria for the 2019 Japanese National Nurse Examination and was close to passing the 2020-2023 examinations, with only a few more correct answers required to pass. ChatGPT had a lower percentage of correct answers in some areas, such as pharmacology, social welfare, related law and regulations, endocrinology/metabolism, and dermatology, and a higher percentage of correct answers in the areas of nutrition, pathology, hematology, ophthalmology, otolaryngology, dentistry and dental surgery, and nursing integration and practice. CONCLUSIONS: ChatGPT only passed the 2019 Japanese National Nursing Examination during the most recent 5 years. Although it did not pass the examinations from other years, it performed very close to the passing level, even in those containing questions related to psychology, communication, and nursing.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,011
score de la tête « metaresearch » (Gemma)0,036
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,011
Score d'incertitude au seuil0,059

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0110,036
Méta-épidémiologie (sens strict)0,0010,000
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0020,001
Études des sciences et des technologies0,0010,001
Communication savante0,0010,002
Science ouverte0,0010,002
Intégrité de la recherche0,0010,001
Charge utile insuffisante (le modèle a refusé de juger)0,0020,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,244
Tête enseignante GPT0,513
Écart entre enseignants0,269 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations79
Publié2023
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueJMIR NursingMême sujetArtificial Intelligence in Healthcare and EducationTravaux en français237 207