MétaCan
Menu
Retour à la cohorte
Enregistrement W4410712214 · doi:10.1016/j.jposna.2025.100196

Artificial Intelligence-Based Large Language Models Can Facilitate Patient Education

2025· article· en· W4410712214 sur OpenAlexaff
Xochitl Bryson, Marleni Albarran, Nicole S. Pham, Arianne Salunga, T Olaf Johnson, Grant D. Hogue, Jaysson T. Brooks, Kali Tileston, Craig R. Louer, Ron El‐Hawary, Meghan N. Imrie, James Policy, Daniel Bouton, Arun R. Hariharan, Sara Van Nortwick, Vidyadhar V. Upasani, Jennifer M. Bauer, Andrew Tice, John S. Vorhies

Notice bibliographique

RevueJournal of the Pediatric Orthopaedic Society of North America · 2025
Typearticle
Langueen
DomaineMedicine
ThématiqueScoliosis diagnosis and treatment
Établissements canadiensChildren's Hospital of Eastern Ontario
Organismes subventionnairesnon disponible
Mots-clésComputer scienceArtificial intelligenceNatural language processingPsychology

Résumé

récupéré en direct d'OpenAlex

Background: Artificial intelligence (AI) large language models (LLMs) are becoming increasingly popular, with patients and families more likely to utilize LLM when conducting internet-based research about scoliosis. For this reason, it is vital to understand the abilities and limitations of this technology in disseminating accurate medical information. We used an expert panel to compare LLM-generated and professional society-authored answers to frequently asked questions about pediatric scoliosis. Methods: We used three publicly available LLMs to generate answers to 15 frequently asked questions (FAQs) regarding pediatric scoliosis. The FAQs were derived from the Scoliosis Research Society, the American Academy of Orthopaedic Surgeons, and the Pediatric Spine Foundation. We gave minimal training to the LLM other than specifying the response length and requesting answers at a 5th-grade reading level. A 15-question survey was distributed to an expert panel composed of pediatric spine surgeons. To determine readability, responses were inputted into an open-source calculator. The panel members were presented with an AI and a physician-generated response to a FAQ and asked to select which they preferred. They were then asked to individually grade the accuracy of responses on a Likert scale. Results: The panel members had a mean of 8.9 years of experience post-fellowship (range: 3-23 years). The panel reported nearly equivalent agreement between AI-generated and physician-generated answers. The expert panel favored professional society-written responses for 40% of questions, AI for 40%, ranked responses equally good for 13%, and saw a tie between AI and "equally good" for 7%. For two professional society-generated and three AI-generated responses, the error bars of the expert panel mean score for accuracy and appropriateness fell below neutral, indicating a lack of consensus and mixed opinions with the response. Conclusions: Based on the expert panel review, AI delivered accurate and appropriate answers as frequently as professional society-authored FAQ answers from professional society websites. AI and professional society websites were equally likely to generate answers with which the expert panel disagreed. Key Concepts: (1)Large language models (LLMs) are increasingly used for generating medical information online, necessitating an evaluation of their accuracy and effectiveness compared with traditional sources.(2)An expert panel of physicians compared artificial intelligence (AI)-generated answers with professional society-authored answers to pediatric scoliosis frequently asked questions, finding that both types of answers were equally favored in terms of accuracy and appropriateness.(3)The panel reported a similar rate of disagreement with AI-generated and professional society-generated answers, indicating that both had areas of controversy.(4)Over half of the expert panel members felt they could distinguish between AI-generated and professional society-generated answers but this did not relate to their preferences.(5)While AI can support medical information dissemination, further research and improvements are needed to address its limitations and ensure high-quality, accessible patient education. Levels of Evidence: IV.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,221
Score d'incertitude au seuil0,440

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,001
Bibliométrie0,0000,001
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,021
Tête enseignante GPT0,279
Écart entre enseignants0,258 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations4
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueJournal of the Pediatric Orthopaedic Society of North AmericaMême sujetScoliosis diagnosis and treatmentTravaux en français237 207