Medical Applications of Artificial Intelligence and Large Language Models: Bibliometric Analysis and Stern Call for Improved Publishing Practices
Bibliographic record
Abstract
The potential medical applications of artificial intelligence (AI) are attracting attention in the scientific literature, with a whopping 222 publications on ChatGPT alone (Figure 1A).1 Large language models such as ChatGPT interpret, synthesize, and output information in the form of text.2 Released by OpenAI (San Francisco, CA) in November 2022, ChatGPT has taken the world by storm, revolutionizing the way physicians interact with AI technology.3 For physicians eager to take advantage of this emerging trend, here we briefly review the available literature describing ChatGPT's accuracy, reliability, and safety. (A) Temporal and (B) geographical publication trends of ChatGPT (OpenAI, San Franciso, CA)-related research in the medical literature. Left y-axis (blue), new publication count; right y-axis (orange), cumulative publication count. Articles in 124 different journals reported on medical or surgical applications of ChatGPT. The publications with the greatest share of ChatGPT articles were Cureus (21.2%; n = 47), Annals of Biomedical Engineering (3.6%; n = 8), and Aesthetic Surgery Journal (2.7%; n = 6). Publications originated in 34 countries, with the highest article volumes coming out of the United States (41.4%; n = 92), China (9.5%; n = 21), and India (6.3%; n = 14) (Figure 1B). These studies have been cited 1354 times at the time of this writing, representing an average of 6.1 citations per paper (range, 0-224). Normalizing to time since publication, ChatGPT articles average 513.2 citations per month, with an average of 2.4 citations per paper per month. These impressive bibliometrics, accumulating in just 6 months since the technology's release, do not necessarily correlate with reliability. Among the 222 publications, 62 articles (27.9%) reported on merely postulated applications of ChatGPT—mostly as letters to the editor. Albeit stimulating for discussion, most of these articles lack evidence and scientific rigor. Of the 121 out of 222 studies (54.5%) reporting on demonstrated applications, 57 out of 121 (47.1%) lacked both validation and ChatGPT performance assessment, thus representing mere “proof of concepts.” Only after AI's performance accuracy, reliability, and safety have been proven will proposed medical applications be usable. Of the 64 studies (52.9%) incorporating some form of performance assessment, the majority of assessments were subjective (34/64, 53.1%), and only 30 out of 64 (46.9%) employed objective validation techniques. Another 39 articles (17.6%) were case reports written with the assistance of ChatGPT, without applicability to medical practice (Figure 2). Publications on ChatGPT (OpenAI, San Franciso, CA) in the medical literature generally lack objective assessment of artificial intelligence performance in its suggested applications. Only 30 out of 222 (13.5%) of articles proposed, implemented, and objectively assessed the performance of ChatGPT across potential medical applications, and the remaining 86.5% of the literature was of little scientific value. Currently, this nascent field falls below scientific publishing standards,4 but with 513.2 citations per month, there is a strong appetite for information on the topic. Commentaries and letters to the editor on theoretical applications serve as inspiration to develop ideas but do not constitute evidence. Reliable research and objective assessment of performance are necessary to bridge the gap between proposed AI applications in medicine and adoption within necessary safety regulations. No components of the present study's conception, design, execution, writing, or editing were done in any part or assisted by ChatGPT. The authors declared no potential conflicts of interest with respect to the research, authorship, and publication of this article. The authors received no financial support for the research, authorship, and publication of this article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.016 | 0.025 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".