Delineating Medical Education: Bibliometric Research Approach(es)
Bibliographic record
Abstract
Abstract Background The field of medical education remains poorly delineated such that there is no broad consensus of the articles and journals that comprise “the field.” This lack of consensus has implications for conducting bibliometric studies and other research designs (e.g., systematic reviews); it also challenges the field to compare citation scores in the field and across others and for an individual to identify themselves as “a medical education researcher.” Other fields have utilized bibliometric field delineation, which is the assigning of articles or journals to a certain field in an effort to define that field. Process In this Research Approach , three bibliometric field delineation approaches -- information retrieval, core journals, and journal co-citation -- are introduced. For each approach, the authors describe their attempt to apply it in the medical education context and identify related strengths and weaknesses. Based on co-citation, the authors propose the Medical Education Journal List 24 (MEJ-24), as a starting point for delineating medical education and invite the community to collaborate on improving and potentially expanding this list. Pearls As a research approach, field delineation is complicated, and there is no clear best way to delineate the field of medical education. However, recent advances in information and computer science provide potentially more fruitful approaches to deal with the complexity of the field. When considering these emerging approaches, researchers should consider collaborating with bibliometricians. Bibliometric approaches rely on available metadata for articles and journals, which necessitates that researchers examine the metadata prior to analysis to understand its strengths and weaknesses, and to assess how this might affect their data interpretation. While using bibliometric approaches for field delineation is valuable, it is important to remember that these techniques are only as good as the research team’s interpretation of the data, which suggests that an expanded research approach is needed to better delineate medical education, an approach that includes active discussion within the medical education community.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.037 | 0.110 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.160 | 0.171 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.017 | 0.011 |
| Open science | 0.002 | 0.007 |
| Research integrity | 0.003 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".