MétaCan
Menu
← Retour à la cohorte
Enregistrement W7116866818 · doi:10.64898/2025.12.19.25341381

A double-blind, crossover, non-inferiority randomized controlled trial where primary care providers and patients compare human- and AI-generated digital health messages: the AI-CARE study protocol

2025· article· W7116866818 sur OpenAlexafffundabout
Audrée Lemieux, S. Kutcher, Borris Rosnay Galani Tietcheu, Gretchen Seitz, Jason Trickovic, Douglas Archibald, Sylvie Grosjean, William Hogg, Sharon Johnston

Notice bibliographique

RevuemedRxiv · 2025
Typearticle
Langue
DomaineMedicine
ThématiqueArtificial Intelligence in Healthcare and Education
Établissements canadiensUniversity of OttawaEastern Ontario Training BoardAgricultural Research Institute of OntarioInstitut du Savoir Montfort
Organismes subventionnairesCanadian Institutes of Health ResearchCHEO Research Institute
Mots-clésProtocol (science)Randomized controlled trialRelevance (law)Primary careCLARITYDigital healthHealth careQuality (philosophy)mHealthPublic health

Résumé

récupéré en direct d'OpenAlex

ABSTRACT Introduction Primary care is facing multiple crises, including an increase in health misinformation. Digital health messaging by primary care providers has been shown to reach a diverse patient population. With the uptake of Generative Artificial Intelligence (GenAI) usage in healthcare, there is an important opportunity to rapidly create messages that are tailored to different populations and conditions. However, thoroughly assessing AI-generated content is essential, as GenAI raises concerns regarding its accuracy, understandability, actionability, and bias perpetuation. We aim to investigate whether digital health messages created by GenAI are evaluated as noninferior compared to those created by human experts. Methods and analysis The AI-CARE (AI to Create Accessible and Reliable patient Education materials) study is a double-blind, crossover, non-inferiority randomized controlled trial. Data collection began on May 30, 2025, and is expected to be completed at the end of April 2026. Over 12 months, 192 messages on 48 topics will be written: half by primary care and public health experts and half by a GenAI tool (OpenAI’s ChatGPT). Review Panels composed of 24 primary care providers and 24 patients will evaluate these messages using an Evaluation Grid developed to assess the messages’ quality of information, adaptation to the target audience, relevance and usefulness, and readiness to be shared with patients. Evaluations will be completed via online REDCap surveys and the order in which the 192 messages appear will be randomized and will vary between individuals. Participants and analysts will be blinded to the generation source. The primary outcome will be the Clarity and Understandability score. Ethics and dissemination The Research Ethics Boards of the Hôpital Montfort (24-25-11-038) and the University of Ottawa (S-12-24-11153) formally approved this study in December 2024. Reported data will be grouped and anonymized for dissemination in peer-reviewed scientific journals and conferences. Trial registration number NCT06997107 ARTICLE SUMMARY Strengths and limitations of this study The AI-CARE study allows for within-participant comparison between human- and AI-generated digital health messages, minimizing variability due to individual differences. The Review Panels are diverse and composed of primary care providers and patients currently practicing in or using the healthcare system in five Canadian provinces. The developed Evaluation Grid allows for the assessment of multiple aspects related to digital health messages: quality of information, adaptability to the target audience, relevance and usefulness, and readiness to be shared with patients. One limitation is that messages generated by AI are created using only one LLM (Open AI’s ChatGPT). Due to the nature and location of recruitment, we may introduce selection bias (participants already engaged in research and interested in digital communication and AI) and the racial diversity of our study population may be limited.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,022
score de la tête « metaresearch » (Gemma)0,025
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Essai randomisé · Signal consensuel: Essai randomisé
GenreSignal candidat: Protocole · Signal consensuel: Protocole
Score de désaccord entre enseignants0,039
Score d'incertitude au seuil0,130

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0220,025
Méta-épidémiologie (sens strict)0,0060,002
Méta-épidémiologie (sens large)0,0090,003
Bibliométrie0,0020,002
Études des sciences et des technologies0,0020,004
Communication savante0,0030,003
Science ouverte0,0030,001
Intégrité de la recherche0,0080,006
Charge utile insuffisante (le modèle a refusé de juger)0,0390,007

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,067
Tête enseignante GPT0,435
Écart entre enseignants0,368 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeEssai randomisé
Domainenon disponible
GenreProtocole

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2025
Routes d'admission3
Résumé présentoui

Explorer davantage

Même revuemedRxiv→Même sujetArtificial Intelligence in Healthcare and Education→Travaux en français237 207→