MétaCan
Menu
Back to cohort
Record W4392691776 · doi:10.1186/s12911-024-02459-6

Assessing the research landscape and clinical utility of large language models: a scoping review

2024· review· en· W4392691776 on OpenAlexaff
Ye‐Jean Park, Abhinav Pillai, Jiawen Deng, Eddie Guo, Mehul Gupta, Mike Paget, Christopher Naugler

Bibliographic record

VenueBMC Medical Informatics and Decision Making · 2024
Typereview
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsCanada Research ChairsUniversity of TorontoUniversity of CalgaryUniversity of New Brunswick
Fundersnot available
KeywordsCINAHLMEDLINEHealth careMedicineSocioeconomic statusPolitical scienceEnvironmental healthPopulation

Abstract

fetched live from OpenAlex

IMPORTANCE: Large language models (LLMs) like OpenAI's ChatGPT are powerful generative systems that rapidly synthesize natural language responses. Research on LLMs has revealed their potential and pitfalls, especially in clinical settings. However, the evolving landscape of LLM research in medicine has left several gaps regarding their evaluation, application, and evidence base. OBJECTIVE: This scoping review aims to (1) summarize current research evidence on the accuracy and efficacy of LLMs in medical applications, (2) discuss the ethical, legal, logistical, and socioeconomic implications of LLM use in clinical settings, (3) explore barriers and facilitators to LLM implementation in healthcare, (4) propose a standardized evaluation framework for assessing LLMs' clinical utility, and (5) identify evidence gaps and propose future research directions for LLMs in clinical applications. EVIDENCE REVIEW: We screened 4,036 records from MEDLINE, EMBASE, CINAHL, medRxiv, bioRxiv, and arXiv from January 2023 (inception of the search) to June 26, 2023 for English-language papers and analyzed findings from 55 worldwide studies. Quality of evidence was reported based on the Oxford Centre for Evidence-based Medicine recommendations. FINDINGS: Our results demonstrate that LLMs show promise in compiling patient notes, assisting patients in navigating the healthcare system, and to some extent, supporting clinical decision-making when combined with human oversight. However, their utilization is limited by biases in training data that may harm patients, the generation of inaccurate but convincing information, and ethical, legal, socioeconomic, and privacy concerns. We also identified a lack of standardized methods for evaluating LLMs' effectiveness and feasibility. CONCLUSIONS AND RELEVANCE: This review thus highlights potential future directions and questions to address these limitations and to further explore LLMs' potential in enhancing healthcare delivery.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.172
metaresearch head score (Gemma)0.548
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Review · Consensus signal: Review
Teacher disagreement score0.828
Threshold uncertainty score0.909

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1720.548
Meta-epidemiology (narrow)0.0020.003
Meta-epidemiology (broad)0.0090.010
Bibliometrics0.0350.025
Science and technology studies0.0020.005
Scholarly communication0.0110.014
Open science0.0050.006
Research integrity0.0060.005
Insufficient payload (model declined to judge)0.0060.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.694
GPT teacher head0.689
Teacher spread0.005 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSystematic review
DomainMethods
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations167
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueBMC Medical Informatics and Decision MakingSame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207