MétaCan
Menu
Back to cohort
Record W4386083791 · doi:10.1093/asj/sjad277

Medical Applications of Artificial Intelligence and Large Language Models: Bibliometric Analysis and Stern Call for Improved Publishing Practices

2023· article· en· W4386083791 on OpenAlexaff
Jad Abi‐Rafeh, Hong Hao Xu, Roy Kazan, Heather Furnas

Bibliographic record

VenueAesthetic Surgery Journal · 2023
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsUniversité LavalMcGill University Health Centre
Fundersnot available
KeywordsMedicineSternPublishingMEDLINEData scienceComputer science

Abstract

fetched live from OpenAlex

The potential medical applications of artificial intelligence (AI) are attracting attention in the scientific literature, with a whopping 222 publications on ChatGPT alone (Figure 1A).1 Large language models such as ChatGPT interpret, synthesize, and output information in the form of text.2 Released by OpenAI (San Francisco, CA) in November 2022, ChatGPT has taken the world by storm, revolutionizing the way physicians interact with AI technology.3 For physicians eager to take advantage of this emerging trend, here we briefly review the available literature describing ChatGPT's accuracy, reliability, and safety. (A) Temporal and (B) geographical publication trends of ChatGPT (OpenAI, San Franciso, CA)-related research in the medical literature. Left y-axis (blue), new publication count; right y-axis (orange), cumulative publication count. Articles in 124 different journals reported on medical or surgical applications of ChatGPT. The publications with the greatest share of ChatGPT articles were Cureus (21.2%; n = 47), Annals of Biomedical Engineering (3.6%; n = 8), and Aesthetic Surgery Journal (2.7%; n = 6). Publications originated in 34 countries, with the highest article volumes coming out of the United States (41.4%; n = 92), China (9.5%; n = 21), and India (6.3%; n = 14) (Figure 1B). These studies have been cited 1354 times at the time of this writing, representing an average of 6.1 citations per paper (range, 0-224). Normalizing to time since publication, ChatGPT articles average 513.2 citations per month, with an average of 2.4 citations per paper per month. These impressive bibliometrics, accumulating in just 6 months since the technology's release, do not necessarily correlate with reliability. Among the 222 publications, 62 articles (27.9%) reported on merely postulated applications of ChatGPT—mostly as letters to the editor. Albeit stimulating for discussion, most of these articles lack evidence and scientific rigor. Of the 121 out of 222 studies (54.5%) reporting on demonstrated applications, 57 out of 121 (47.1%) lacked both validation and ChatGPT performance assessment, thus representing mere “proof of concepts.” Only after AI's performance accuracy, reliability, and safety have been proven will proposed medical applications be usable. Of the 64 studies (52.9%) incorporating some form of performance assessment, the majority of assessments were subjective (34/64, 53.1%), and only 30 out of 64 (46.9%) employed objective validation techniques. Another 39 articles (17.6%) were case reports written with the assistance of ChatGPT, without applicability to medical practice (Figure 2). Publications on ChatGPT (OpenAI, San Franciso, CA) in the medical literature generally lack objective assessment of artificial intelligence performance in its suggested applications. Only 30 out of 222 (13.5%) of articles proposed, implemented, and objectively assessed the performance of ChatGPT across potential medical applications, and the remaining 86.5% of the literature was of little scientific value. Currently, this nascent field falls below scientific publishing standards,4 but with 513.2 citations per month, there is a strong appetite for information on the topic. Commentaries and letters to the editor on theoretical applications serve as inspiration to develop ideas but do not constitute evidence. Reliable research and objective assessment of performance are necessary to bridge the gap between proposed AI applications in medicine and adoption within necessary safety regulations. No components of the present study's conception, design, execution, writing, or editing were done in any part or assisted by ChatGPT. The authors declared no potential conflicts of interest with respect to the research, authorship, and publication of this article. The authors received no financial support for the research, authorship, and publication of this article.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.093
metaresearch head score (Gemma)0.262
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Bibliometrics
Consensus categoriesnone
DomainCandidate signal: Reporting · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.932
Threshold uncertainty score0.492

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0930.262
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0030.002
Bibliometrics0.0680.119
Science and technology studies0.0010.003
Scholarly communication0.0130.015
Open science0.0020.004
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.208
GPT teacher head0.441
Teacher spread0.234 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainReporting
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations8
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueAesthetic Surgery JournalSame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207