MétaCan
Menu
Back to cohort
Record W4415788501 · doi:10.1186/s12891-025-09227-1

Evaluating the performance of five large language models in answering Delphi consensus questions relating to patellar instability and medial patellofemoral ligament reconstruction

2025· article· en· W4415788501 on OpenAlexaff
Prushoth Vivekanantha, Dan Cohen, David Slawaska‐Eng, Kanto Nagai, Magdalena Tarchala, Bogdan A. Matache, Laurie A. Hiemstra, Robert Longstaffe, Bryson P. Lesniak, Amit Meena, Sachin Tapasvi, Petri Sillanpää, Patrick Grzela, Daniel L. Lamanna, Kristian Samuelsson, Darren de

Bibliographic record

VenueBMC Musculoskeletal Disorders · 2025
Typearticle
Languageen
FieldMedicine
TopicTotal Knee Arthroplasty Outcomes
Canadian institutionsUniversity of ManitobaUniversity of CalgaryUniversity of OttawaMcMaster University
FundersGöteborgs Universitet
KeywordsMedial patellofemoral ligamentSports medicineDelphi methodOrthopedic surgeryRehabilitationPatellofemoral pain syndromeMultidisciplinary approachDelphi

Abstract

fetched live from OpenAlex

PURPOSE: Artificial intelligence (AI) has become incredibly popular over the past several years, with large language models (LLMs) offering the possibility of revolutionizing the way healthcare information is shared with patients. However, to prevent the spread of misinformation, analyzing the accuracy of answers from these LLMs is essential. This study will aim to assess the accuracy of five freely accessible chatbots by specifically evaluating their responses to questions about patellofemoral instability (PFI). The secondary objective will be to compare the different chatbots, to distinguish which LLM offers the most accurate set of responses. METHODS: Ten questions were selected from a previously published international Delphi Consensus study pertaining to patellar instability, and posed to ChatGPT4o, Perplexity AI, Bing CoPilot, Claude2, and Google Gemini. Responses were assessed for accuracy using the validated Mika score by eight Orthopedic surgeons who have completed fellowship training in sports-medicine. Median responses amongst the eight reviewers for each question were compared using the Kruskal-Wallis and Dunn's post-hoc tests. Percentages of each Mika score distribution were compared using Pearson's chi-square test. P-values less than or equal to 0.05 were considered significant. The Gwet's AC2 coefficient was calculated to assess for inter-rater agreement, corrected for chance and employing quadratic weights. RESULTS: ChatGPT4o and Claude2 had the highest percentage of reviews (38/80, 47.5%) considered to be an "excellent response not requiring classification", or a Mika score of 1. Google Gemini had the highest percentage of reviews (17/80, 21.3%) considered to be "unsatisfactory requiring substantial clarification", or a Mika score of 4 (p < 0.001). The median ± interquartile range (IQR) Mika scores was 2 (1) for ChatGPT4o and Perplexity AI, 2 (2) for Bing CoPilot and Claude2, and 3 (2) for Google Gemini. Median responses were not significantly different between ChatGPT4o, Perplexity AI, Bing CoPilot, and Claude2, however all four statistically outperformed Google Gemini (p < 0.05). Inter-rater agreement was classified as moderate (0.40 > AC2 ≥ 0.60) for ChatGPT, Perplexity AI, Bing CoPilot, and Claude2, while there was no agreement for Google Gemini (AC2 < 0). CONCLUSION: Current free access LLMs (ChatGPT4o, Perplexity AI, Bing CoPilot, and Claude2) predominantly provide satisfactory responses requiring minimal clarification to standardized questions relating to patellar instability. Google Gemini statistically underperformed in accuracy relative to the other four LLMs, with most answers requiring moderate clarification. Furthermore, inter-rater agreement was moderate for all LLMs apart from Google Gemini, which had no agreement. These findings advocate for the utility of existing LLMs in serving as an adjunct to physicians and surgeons in providing patients information pertaining to patellar instability. LEVEL OF EVIDENCE: V.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.238
Threshold uncertainty score0.559

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.021
GPT teacher head0.322
Teacher spread0.301 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueBMC Musculoskeletal DisordersSame topicTotal Knee Arthroplasty OutcomesFrench-language works237,207