MétaCan
Menu
Retour à la cohorte
Enregistrement W4386083791 · doi:10.1093/asj/sjad277

Medical Applications of Artificial Intelligence and Large Language Models: Bibliometric Analysis and Stern Call for Improved Publishing Practices

2023· article· en· W4386083791 sur OpenAlexaff
Jad Abi‐Rafeh, Hong Hao Xu, Roy Kazan, Heather Furnas

Notice bibliographique

RevueAesthetic Surgery Journal · 2023
Typearticle
Langueen
DomaineMedicine
ThématiqueArtificial Intelligence in Healthcare and Education
Établissements canadiensUniversité LavalMcGill University Health Centre
Organismes subventionnairesnon disponible
Mots-clésMedicineSternPublishingMEDLINEData scienceComputer science

Résumé

récupéré en direct d'OpenAlex

The potential medical applications of artificial intelligence (AI) are attracting attention in the scientific literature, with a whopping 222 publications on ChatGPT alone (Figure 1A).1 Large language models such as ChatGPT interpret, synthesize, and output information in the form of text.2 Released by OpenAI (San Francisco, CA) in November 2022, ChatGPT has taken the world by storm, revolutionizing the way physicians interact with AI technology.3 For physicians eager to take advantage of this emerging trend, here we briefly review the available literature describing ChatGPT's accuracy, reliability, and safety. (A) Temporal and (B) geographical publication trends of ChatGPT (OpenAI, San Franciso, CA)-related research in the medical literature. Left y-axis (blue), new publication count; right y-axis (orange), cumulative publication count. Articles in 124 different journals reported on medical or surgical applications of ChatGPT. The publications with the greatest share of ChatGPT articles were Cureus (21.2%; n = 47), Annals of Biomedical Engineering (3.6%; n = 8), and Aesthetic Surgery Journal (2.7%; n = 6). Publications originated in 34 countries, with the highest article volumes coming out of the United States (41.4%; n = 92), China (9.5%; n = 21), and India (6.3%; n = 14) (Figure 1B). These studies have been cited 1354 times at the time of this writing, representing an average of 6.1 citations per paper (range, 0-224). Normalizing to time since publication, ChatGPT articles average 513.2 citations per month, with an average of 2.4 citations per paper per month. These impressive bibliometrics, accumulating in just 6 months since the technology's release, do not necessarily correlate with reliability. Among the 222 publications, 62 articles (27.9%) reported on merely postulated applications of ChatGPT—mostly as letters to the editor. Albeit stimulating for discussion, most of these articles lack evidence and scientific rigor. Of the 121 out of 222 studies (54.5%) reporting on demonstrated applications, 57 out of 121 (47.1%) lacked both validation and ChatGPT performance assessment, thus representing mere “proof of concepts.” Only after AI's performance accuracy, reliability, and safety have been proven will proposed medical applications be usable. Of the 64 studies (52.9%) incorporating some form of performance assessment, the majority of assessments were subjective (34/64, 53.1%), and only 30 out of 64 (46.9%) employed objective validation techniques. Another 39 articles (17.6%) were case reports written with the assistance of ChatGPT, without applicability to medical practice (Figure 2). Publications on ChatGPT (OpenAI, San Franciso, CA) in the medical literature generally lack objective assessment of artificial intelligence performance in its suggested applications. Only 30 out of 222 (13.5%) of articles proposed, implemented, and objectively assessed the performance of ChatGPT across potential medical applications, and the remaining 86.5% of the literature was of little scientific value. Currently, this nascent field falls below scientific publishing standards,4 but with 513.2 citations per month, there is a strong appetite for information on the topic. Commentaries and letters to the editor on theoretical applications serve as inspiration to develop ideas but do not constitute evidence. Reliable research and objective assessment of performance are necessary to bridge the gap between proposed AI applications in medicine and adoption within necessary safety regulations. No components of the present study's conception, design, execution, writing, or editing were done in any part or assisted by ChatGPT. The authors declared no potential conflicts of interest with respect to the research, authorship, and publication of this article. The authors received no financial support for the research, authorship, and publication of this article.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,004
score de la tête « metaresearch » (Gemma)0,002
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesBibliométrie
Catégories consensuellesBibliométrie
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Autre devis · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,981
Score d'incertitude au seuil0,996

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0040,002
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0160,025
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,208
Tête enseignante GPT0,441
Écart entre enseignants0,234 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.

Devis d'étudeAutre devis
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations8
Publié2023
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueAesthetic Surgery JournalMême sujetArtificial Intelligence in Healthcare and EducationTravaux en français237 207