MétaCan
Menu
Back to cohort
Record W4386083791 · doi:10.1093/asj/sjad277

Medical Applications of Artificial Intelligence and Large Language Models: Bibliometric Analysis and Stern Call for Improved Publishing Practices

2023· article· en· W4386083791 on OpenAlexaff
Jad Abi‐Rafeh, Hong Hao Xu, Roy Kazan, Heather Furnas

Bibliographic record

VenueAesthetic Surgery Journal · 2023
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsUniversité LavalMcGill University Health Centre
Fundersnot available
KeywordsMedicineSternPublishingMEDLINEData scienceComputer science

Abstract

fetched live from OpenAlex

The potential medical applications of artificial intelligence (AI) are attracting attention in the scientific literature, with a whopping 222 publications on ChatGPT alone (Figure 1A).1 Large language models such as ChatGPT interpret, synthesize, and output information in the form of text.2 Released by OpenAI (San Francisco, CA) in November 2022, ChatGPT has taken the world by storm, revolutionizing the way physicians interact with AI technology.3 For physicians eager to take advantage of this emerging trend, here we briefly review the available literature describing ChatGPT's accuracy, reliability, and safety. (A) Temporal and (B) geographical publication trends of ChatGPT (OpenAI, San Franciso, CA)-related research in the medical literature. Left y-axis (blue), new publication count; right y-axis (orange), cumulative publication count. Articles in 124 different journals reported on medical or surgical applications of ChatGPT. The publications with the greatest share of ChatGPT articles were Cureus (21.2%; n = 47), Annals of Biomedical Engineering (3.6%; n = 8), and Aesthetic Surgery Journal (2.7%; n = 6). Publications originated in 34 countries, with the highest article volumes coming out of the United States (41.4%; n = 92), China (9.5%; n = 21), and India (6.3%; n = 14) (Figure 1B). These studies have been cited 1354 times at the time of this writing, representing an average of 6.1 citations per paper (range, 0-224). Normalizing to time since publication, ChatGPT articles average 513.2 citations per month, with an average of 2.4 citations per paper per month. These impressive bibliometrics, accumulating in just 6 months since the technology's release, do not necessarily correlate with reliability. Among the 222 publications, 62 articles (27.9%) reported on merely postulated applications of ChatGPT—mostly as letters to the editor. Albeit stimulating for discussion, most of these articles lack evidence and scientific rigor. Of the 121 out of 222 studies (54.5%) reporting on demonstrated applications, 57 out of 121 (47.1%) lacked both validation and ChatGPT performance assessment, thus representing mere “proof of concepts.” Only after AI's performance accuracy, reliability, and safety have been proven will proposed medical applications be usable. Of the 64 studies (52.9%) incorporating some form of performance assessment, the majority of assessments were subjective (34/64, 53.1%), and only 30 out of 64 (46.9%) employed objective validation techniques. Another 39 articles (17.6%) were case reports written with the assistance of ChatGPT, without applicability to medical practice (Figure 2). Publications on ChatGPT (OpenAI, San Franciso, CA) in the medical literature generally lack objective assessment of artificial intelligence performance in its suggested applications. Only 30 out of 222 (13.5%) of articles proposed, implemented, and objectively assessed the performance of ChatGPT across potential medical applications, and the remaining 86.5% of the literature was of little scientific value. Currently, this nascent field falls below scientific publishing standards,4 but with 513.2 citations per month, there is a strong appetite for information on the topic. Commentaries and letters to the editor on theoretical applications serve as inspiration to develop ideas but do not constitute evidence. Reliable research and objective assessment of performance are necessary to bridge the gap between proposed AI applications in medicine and adoption within necessary safety regulations. No components of the present study's conception, design, execution, writing, or editing were done in any part or assisted by ChatGPT. The authors declared no potential conflicts of interest with respect to the research, authorship, and publication of this article. The authors received no financial support for the research, authorship, and publication of this article.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesBibliometrics
Consensus categoriesBibliometrics
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.981
Threshold uncertainty score0.996

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0040.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0160.025
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.208
GPT teacher head0.441
Teacher spread0.234 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations8
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueAesthetic Surgery JournalSame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207