Medical Applications of Artificial Intelligence and Large Language Models: Bibliometric Analysis and Stern Call for Improved Publishing Practices
Notice bibliographique
Résumé
The potential medical applications of artificial intelligence (AI) are attracting attention in the scientific literature, with a whopping 222 publications on ChatGPT alone (Figure 1A).1 Large language models such as ChatGPT interpret, synthesize, and output information in the form of text.2 Released by OpenAI (San Francisco, CA) in November 2022, ChatGPT has taken the world by storm, revolutionizing the way physicians interact with AI technology.3 For physicians eager to take advantage of this emerging trend, here we briefly review the available literature describing ChatGPT's accuracy, reliability, and safety. (A) Temporal and (B) geographical publication trends of ChatGPT (OpenAI, San Franciso, CA)-related research in the medical literature. Left y-axis (blue), new publication count; right y-axis (orange), cumulative publication count. Articles in 124 different journals reported on medical or surgical applications of ChatGPT. The publications with the greatest share of ChatGPT articles were Cureus (21.2%; n = 47), Annals of Biomedical Engineering (3.6%; n = 8), and Aesthetic Surgery Journal (2.7%; n = 6). Publications originated in 34 countries, with the highest article volumes coming out of the United States (41.4%; n = 92), China (9.5%; n = 21), and India (6.3%; n = 14) (Figure 1B). These studies have been cited 1354 times at the time of this writing, representing an average of 6.1 citations per paper (range, 0-224). Normalizing to time since publication, ChatGPT articles average 513.2 citations per month, with an average of 2.4 citations per paper per month. These impressive bibliometrics, accumulating in just 6 months since the technology's release, do not necessarily correlate with reliability. Among the 222 publications, 62 articles (27.9%) reported on merely postulated applications of ChatGPT—mostly as letters to the editor. Albeit stimulating for discussion, most of these articles lack evidence and scientific rigor. Of the 121 out of 222 studies (54.5%) reporting on demonstrated applications, 57 out of 121 (47.1%) lacked both validation and ChatGPT performance assessment, thus representing mere “proof of concepts.” Only after AI's performance accuracy, reliability, and safety have been proven will proposed medical applications be usable. Of the 64 studies (52.9%) incorporating some form of performance assessment, the majority of assessments were subjective (34/64, 53.1%), and only 30 out of 64 (46.9%) employed objective validation techniques. Another 39 articles (17.6%) were case reports written with the assistance of ChatGPT, without applicability to medical practice (Figure 2). Publications on ChatGPT (OpenAI, San Franciso, CA) in the medical literature generally lack objective assessment of artificial intelligence performance in its suggested applications. Only 30 out of 222 (13.5%) of articles proposed, implemented, and objectively assessed the performance of ChatGPT across potential medical applications, and the remaining 86.5% of the literature was of little scientific value. Currently, this nascent field falls below scientific publishing standards,4 but with 513.2 citations per month, there is a strong appetite for information on the topic. Commentaries and letters to the editor on theoretical applications serve as inspiration to develop ideas but do not constitute evidence. Reliable research and objective assessment of performance are necessary to bridge the gap between proposed AI applications in medicine and adoption within necessary safety regulations. No components of the present study's conception, design, execution, writing, or editing were done in any part or assisted by ChatGPT. The authors declared no potential conflicts of interest with respect to the research, authorship, and publication of this article. The authors received no financial support for the research, authorship, and publication of this article.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,016 | 0,025 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».