A Systematic Framework for Analyzing Patient-Generated Narrative Data: Protocol for a Content Analysis
Notice bibliographique
Résumé
BACKGROUND: Patient narrative data in online health care forums (communities) are receiving increasing attention from the scientific community for implementing patient-centered care. Natural language processing (NLP) methods are gaining more and more attention because of the enormous data volume. However, state-of-the-art NLP still cannot meet the need of high-resolution analysis of patients' narratives. Manual qualitative analysis still plays a pivotal role in answering complicated research questions from analyzing patient narratives. OBJECTIVE: This study aimed to develop a systematic framework for qualitative analysis of patient-generated narratives in online health care forums. METHODS: Our systematic framework consists of 4 phases: (1) data collection, (2) data preparation, (3) content analysis, and (4) interpretation of the results. Data collection and data preparation phases are constructed based on text mining methods for identifying appropriate online health forums for data collection, differentiating posts of patients from other stakeholders, protecting patients' privacy, sampling, and choosing the unit of analysis. Content analysis phase is built on the framework method, which facilitates and accelerates the identification of patterns and themes by an interdisciplinary research team. In the end, the focus of interpretation of the results phase is to measure the data quality and interpret the findings regarding the dimensions and aspects of patients' experiences and concerns in their original contexts. RESULTS: We demonstrated the usability of the proposed systematic framework using 2 case studies: one on determining factors affecting patients' attitudes toward antidepressants and another on identifying the disease management strategies in patient with diabetes facing financial difficulties. The framework provides a clear step-by-step process for systematic content analysis of patient narratives and produces high-quality structured results that can be used for describing patterns or regularities in patients' experiences, generating and testing hypotheses, and identifying areas of improvement in the health care systems. CONCLUSIONS: The systematic framework is a rigorous and standardized method for qualitative analysis of patient narratives. Findings obtained through such a process indicate authentic dimensions and aspects of patient experiences and shed light on patients' concerns, needs, preferences, and values, which are the core of patient-centered care. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): RR1-10.2196/13914.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,201 | 0,198 |
| Méta-épidémiologie (sens strict) | 0,004 | 0,005 |
| Méta-épidémiologie (sens large) | 0,005 | 0,006 |
| Bibliométrie | 0,013 | 0,011 |
| Études des sciences et des technologies | 0,008 | 0,007 |
| Communication savante | 0,006 | 0,005 |
| Science ouverte | 0,005 | 0,006 |
| Intégrité de la recherche | 0,007 | 0,008 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,042 | 0,010 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».