MétaCan
Menu
Back to cohort
Record W4415788501 · doi:10.1186/s12891-025-09227-1

Evaluating the performance of five large language models in answering Delphi consensus questions relating to patellar instability and medial patellofemoral ligament reconstruction

2025· article· en· W4415788501 on OpenAlexaff
Prushoth Vivekanantha, Dan Cohen, David Slawaska‐Eng, Kanto Nagai, Magdalena Tarchala, Bogdan A. Matache, Laurie A. Hiemstra, Robert Longstaffe, Bryson P. Lesniak, Amit Meena, Sachin Tapasvi, Petri Sillanpää, Patrick Grzela, Daniel L. Lamanna, Kristian Samuelsson, Darren de

Bibliographic record

VenueBMC Musculoskeletal Disorders · 2025
Typearticle
Languageen
FieldMedicine
TopicTotal Knee Arthroplasty Outcomes
Canadian institutionsUniversity of ManitobaUniversity of CalgaryUniversity of OttawaMcMaster University
FundersGöteborgs Universitet
KeywordsMedial patellofemoral ligamentSports medicineDelphi methodOrthopedic surgeryRehabilitationPatellofemoral pain syndromeMultidisciplinary approachDelphi

Abstract

fetched live from OpenAlex

PURPOSE: Artificial intelligence (AI) has become incredibly popular over the past several years, with large language models (LLMs) offering the possibility of revolutionizing the way healthcare information is shared with patients. However, to prevent the spread of misinformation, analyzing the accuracy of answers from these LLMs is essential. This study will aim to assess the accuracy of five freely accessible chatbots by specifically evaluating their responses to questions about patellofemoral instability (PFI). The secondary objective will be to compare the different chatbots, to distinguish which LLM offers the most accurate set of responses. METHODS: Ten questions were selected from a previously published international Delphi Consensus study pertaining to patellar instability, and posed to ChatGPT4o, Perplexity AI, Bing CoPilot, Claude2, and Google Gemini. Responses were assessed for accuracy using the validated Mika score by eight Orthopedic surgeons who have completed fellowship training in sports-medicine. Median responses amongst the eight reviewers for each question were compared using the Kruskal-Wallis and Dunn's post-hoc tests. Percentages of each Mika score distribution were compared using Pearson's chi-square test. P-values less than or equal to 0.05 were considered significant. The Gwet's AC2 coefficient was calculated to assess for inter-rater agreement, corrected for chance and employing quadratic weights. RESULTS: ChatGPT4o and Claude2 had the highest percentage of reviews (38/80, 47.5%) considered to be an "excellent response not requiring classification", or a Mika score of 1. Google Gemini had the highest percentage of reviews (17/80, 21.3%) considered to be "unsatisfactory requiring substantial clarification", or a Mika score of 4 (p < 0.001). The median ± interquartile range (IQR) Mika scores was 2 (1) for ChatGPT4o and Perplexity AI, 2 (2) for Bing CoPilot and Claude2, and 3 (2) for Google Gemini. Median responses were not significantly different between ChatGPT4o, Perplexity AI, Bing CoPilot, and Claude2, however all four statistically outperformed Google Gemini (p < 0.05). Inter-rater agreement was classified as moderate (0.40 > AC2 ≥ 0.60) for ChatGPT, Perplexity AI, Bing CoPilot, and Claude2, while there was no agreement for Google Gemini (AC2 < 0). CONCLUSION: Current free access LLMs (ChatGPT4o, Perplexity AI, Bing CoPilot, and Claude2) predominantly provide satisfactory responses requiring minimal clarification to standardized questions relating to patellar instability. Google Gemini statistically underperformed in accuracy relative to the other four LLMs, with most answers requiring moderate clarification. Furthermore, inter-rater agreement was moderate for all LLMs apart from Google Gemini, which had no agreement. These findings advocate for the utility of existing LLMs in serving as an adjunct to physicians and surgeons in providing patients information pertaining to patellar instability. LEVEL OF EVIDENCE: V.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.096
metaresearch head score (Gemma)0.239
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.904
Threshold uncertainty score0.506

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0960.239
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.001
Science and technology studies0.0020.002
Scholarly communication0.0020.002
Open science0.0010.005
Research integrity0.0020.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.021
GPT teacher head0.322
Teacher spread0.301 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSimulation or modeling
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueBMC Musculoskeletal DisordersSame topicTotal Knee Arthroplasty OutcomesFrench-language works237,207