MétaCan
Menu
← Back to cohort
Record W7114920428 · doi:10.2196/69707

Using Large Language Models to Summarize Evidence in Biomedical Articles: Exploratory Comparison Between AI- and Human-Annotated Bibliographies

2025· article· en· W7114920428 on OpenAlexvenueno aff

Bibliographic record

VenueJMIR Formative Research · 2025
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsnot available
Fundersnot available
KeywordsChatbotContext (archaeology)Task (project management)Quality (philosophy)Language modelExploratory research

Abstract

fetched live from OpenAlex

Background: Annotated bibliographies summarize literature, but training, experience, and time are needed to create concise yet accurate annotations. Summaries generated by artificial intelligence (AI) can save human resources, but AI-generated content can also contain serious errors. Objective: To determine the feasibility of using AI as an alternative to human annotators, we explored whether ChatGPT can generate annotations with characteristics that are comparable to those written by humans. Methods: We had 2 humans and 3 versions of ChatGPT (3.5, 4, and 5) independently write annotations on the same set of 15 publications. We collected data on word count and Flesch Reading Ease (FRE). In this study, 2 assessors who were masked to the source of the annotations independently evaluated (1) capture of main points, (2) presence of errors, and (3) whether the annotation included a discussion of both the quality and context of the article within the broader literature. We evaluated agreement and disagreement between the assessors and used descriptive statistics and assessor-stratified binary and cumulative mixed-effects logit models to compare annotations written by ChatGPT and humans. Results: On average, humans wrote shorter annotations (mean 90.20, SD 36.8 words) than ChatGPT (mean 113, SD 16 words) which were easier to interpret (human FRE score, mean 15.3, SD 12.4; ChatGPT FRE score, mean 5.76, SD 7.32). Our assessments of agreement and disagreement revealed that one assessor was consistently stricter than the other. However, assessor-stratified models of main points, errors, and quality/context showed similar qualitative conclusions. There was no statistically significant difference in the odds of presenting a better summary of main points between ChatGPT- and human-generated annotations for either assessor (Assessor 1: OR 0.96, 95% CI 0.12-7.71; Assessor 2: OR 1.64, 95% CI 0.67-4.06). However, both assessors observed that human annotations had lower odds of having one or more types of errors compared to ChatGPT (Assessor 1: OR 0.31, 95% CI 0.09-1.02; Assessor 2: OR 0.10, 95% CI 0.03-0.33). On the other hand, human annotations also had lower odds of summarizing the paper's quality and context when compared to ChatGPT (Assessor 1: OR 0.11, 95% CI 0.03-0.33; Assessor 2: OR 0.03, 95% CI 0.01-0.10). That said, ChatGPT's summaries of quality and context were sometimes inaccurate. Conclusions: Rapidly learning a body of scientific literature is a vital yet daunting task that may be made more efficient by AI tools. In our study, ChatGPT quickly generated concise summaries of academic literature and also provided quality and context more consistently than humans. However, ChatGPT's discussion of the quality and context was not always accurate, and ChatGPT annotations included more errors. Annotated bibliographies that are AI-generated and carefully verified by humans may thus be an efficient way to provide a rapid overview of literature. More research is needed to determine the extent that prompt engineering can reduce errors and improve chatbot performance.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.149
metaresearch head score (Gemma)0.521
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.851
Threshold uncertainty score0.789

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1490.521
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.003
Bibliometrics0.0230.015
Science and technology studies0.0020.003
Scholarly communication0.0080.008
Open science0.0020.006
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.564
GPT teacher head0.609
Teacher spread0.045 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Formative Research→Same topicArtificial Intelligence in Healthcare and Education→French-language works237,207→