Automated de-identification and unstructured textual electronic medical record data in Manitoba
Notice bibliographique
Résumé
Introduction: Unstructured textual electronic medical record (EMR) data contain valuable patient details that can benefit health research. Personal health information (PHI) must be de-identified for EMR data to be used for secondary purposes. A considerable amount of de-identification research has been conducted using existing synthetic, de-identified, and annotated data sets. To date, little is known about how existing de-identification literature applies to unstructured EMR data in Manitoba. Objectives: The research objectives were to: 1) categorize the types and frequency of PHI in Manitoba EMR data, 2) assess the applicability of de-identification literature on Manitoba EMR data, and 3) test how NLM-Scrubber, an existing de-identification tool validated to be successful, redacts PHI in Manitoba EMR data. Methods: The Manitoba data set comprised of 750 unstructured textual EMR encounter notes from 2003 to 2017 from the Manitoba Primary Care Research Network. In-scope PHI included name, personal health information number, address, phone number, and date (excluding year). Two annotators tagged PHI in the Manitoba data using the Visual Tagging Tool. Comparison of Manitoba data and the 2014 i2b2 corpus examined note compilation and PHI prevalence. NLM-Scrubber’s de-identification of Manitoba data was assessed using performance measures and tested against the null hypothesis that NLM-Scrubber will recall ≥87% of PHI in Manitoba data. Results: The Manitoba EMR data contained 3,314 PHI instances, demonstrating 1.6% PHI prevalence. All in-scope PHI types were present. The Manitoba data offered more independent notes and broader variety of note types than the i2b2 corpus. The Manitoba EMR data contained nearly twice as many name PHI instances as the i2b2 corpus (62% and 32%, respectively) but fewer date instances (31% and 55%, respectively). NLM-Scrubber’s PHI recall was 75.4% (95% CI, 72.9-77.8%), leading to rejection of the null hypothesis. Conclusion: Direct and indirect PHI represent a small proportion of Manitoba EMR data. De-identification literature may have limited applicability to Manitoba EMR data. NLM-Scrubber may not be acceptable for use on Manitoba EMR data due to its low recall performance. Attention should be directed to trained machine learning solutions that enable customization, adjustment of rule-based methods, and pseudo PHI to protect patient privacy.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,025 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,006 | 0,008 |
| Études des sciences et des technologies | 0,002 | 0,001 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,001 | 0,004 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».