Identification of Key Nodes and Global Research Trends of artificial intelligence (AI)/ Large Language Model (LLM) in Medical Education: A Bibliometric Analysis(1986-2024) (Preprint)
Bibliographic record
Abstract
BACKGROUND In recent years, artificial intelligence (AI)/ Large Language Model(LLM) has significantly transformed the field of medical education, prompting extensive research. Many important bibliometric nodes in this emerging field have yet to be explored. OBJECTIVE This study aims to synthesize diverse publications to analyze the bibliometric attributes of this field, including landmarks, emerging topics, and development status. Also, the formation pattern of bibliometrics in emerging fields will be analyzed simultaneously. METHODS We utilized the Web of Science Core Collection to download literature on AI in medical education. Bibliometric analysis was performed using Citespace v.6.3.R1 software, which facilitated the analysis of publication volume, collaboration within the field, citation networks, and keyword analysis. Additionally, we employed the Bibliometrix package based on R for generating conceptual and thematic maps related to the topic. RESULTS A total of 547 publications were retrieved from the Web of Science Core Collection, covering the period from 1986 to 2024. The five leading countries in terms of publication volume were the United States, England, China, Canada, and India. The most prolific journals included JMIR Medical Education, BMC Medical Education, Cureus Journal of Medical Science, Medical Teacher, and Academic Medicine. The top institutions contributing to this body of work were the University of London, National University of Singapore, Harvard University, and Stanford University. Other important bibliometric characteristics, such as high-yield authors, highly cited authors, and frequently collaborating authors, were also identified. A citation co-citation network was established to determine the key knowledge base and potential pivotal literature in the field. Citespace software was utilized to identify clusters and bursts of high-frequency terms, highlighting current hotspots within the discipline. The Bibliometrix toolkit provided conceptual and thematic maps to assess the development status and trends in the field. CONCLUSIONS AI/LLM in medical education has emerged as a burgeoning field in recent years. JMIR Medical Education was identified as a key node based on its notable bibliometric characteristics. In the early stages of this emerging discipline, the journal's submission calls significantly influence the bibliometric features of the literature, thereby promoting field development. The discipline is currently in a developmental phase, lacking well-defined subfields. Topics such as "nursing education," "digital health," "medical exams," and "conversational agents" have garnered increasing interest over time. Research related to ChatGPT and large language models appears to occupy a central and influential position. Furthermore, medical ethics, medical training, and skills training are emerging focuses of current development and innovation, particularly in gene technology. However, this analysis indicates that there has been insufficient attention given to clinical reasoning, undergraduate education, and virtual reality in the context of AI/LLM in medical education.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.022 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.049 | 0.091 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".