MétaCan
Menu
Retour à la cohorte
Enregistrement W4413188097 · doi:10.2196/76681

Identifying Firearm Violence Exposure in Primary Care Clinical Notes: Protocol for Developing a National Language Processing Text Classifier

2025· article· en· W4413188097 sur OpenAlexvenueno aff
Natalie A. Cartwright, Frances M. Biel, Megan Hoopes, Ali Al Bataineh, Pedro Rivera, Kerry Ann Bet, Nicole Cook

Notice bibliographique

RevueJMIR Research Protocols · 2025
Typearticle
Langueen
DomaineSocial Sciences
ThématiqueGun Ownership and Violence Research
Établissements canadiensnon disponible
Organismes subventionnairesNational Institutes of Health
Mots-clésPreprintProtocol (science)Computer scienceMedical emergencyMedicineComputer securityNatural language processingWorld Wide WebAlternative medicinePathology

Résumé

récupéré en direct d'OpenAlex

BACKGROUND: Structured data codes capture acute bodily injury from firearm violence but do not necessarily describe follow-up care from bodily injury and secondary exposure to firearm violence (eg, witnessing a shooting, being threatened by a firearm, or losing a loved one to gun violence and injury from firearms) even though such exposure is associated with many short- and long-term health impacts. Clinical notes from electronic health records (EHRs) often contain data not otherwise captured in structured data fields and can be categorized using natural language processing (NLP). OBJECTIVE: This study protocol outlines the steps being taken to develop an NLP text classifier for determination of exposure to firearm violence (both primary and secondary exposure) from ambulatory primary care and behavioral health EHR clinical notes for persons aged ≥5 years. METHODS: The study will use unstructured data from clinical notes taken between 2012 and 2022 from OCHIN, a multistate network of community health organizations using a single instance of Epic EHR. We describe the process of developing a labeled dataset for supervised NLP development that includes establishing a lexicon (words related to firearm violence) to identify potentially relevant notes, followed by a review of text extracted from a sample of these notes. We then describe the process of building, training, and evaluating candidate machine learning, neural network, and large language model NLP text classifiers. From this, a final NLP model is chosen then evaluated on a new set of randomly selected notes. An engaged stakeholder advisory committee will provide input and guidance on methods and results to identify and address potential biases in the NLP text classifiers. RESULTS: The study was funded in September 2023. Study activities have been ongoing through July 2025 and we are currently evaluating NLP text classifiers. We expect that the final model will be selected by August 2025 and we will publish results of NLP model development and the final model performance in 2026. CONCLUSIONS: This work describes the development of a novel NLP text classifier to identify exposure to firearm violence in ambulatory primary care and behavioral health clinical notes. The NLP model developed in this study may lead to increased ascertainment of patients with exposure, laying the groundwork for understanding the long-term impacts and outcomes of firearm violence exposure and presenting opportunities for improved patient care. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/76681.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,039
score de la tête « metaresearch » (Gemma)0,047
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Protocole · Signal consensuel: Protocole
Score de désaccord entre enseignants0,048
Score d'incertitude au seuil0,207

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0390,047
Méta-épidémiologie (sens strict)0,0020,002
Méta-épidémiologie (sens large)0,0020,002
Bibliométrie0,0020,002
Études des sciences et des technologies0,0030,002
Communication savante0,0020,002
Science ouverte0,0030,002
Intégrité de la recherche0,0020,004
Charge utile insuffisante (le modèle a refusé de juger)0,0480,016

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,517
Tête enseignante GPT0,676
Écart entre enseignants0,158 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreProtocole

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueJMIR Research ProtocolsMême sujetGun Ownership and Violence ResearchTravaux en français237 207