Identifying Firearm Violence Exposure in Primary Care Clinical Notes: Protocol for Developing a National Language Processing Text Classifier
Bibliographic record
Abstract
BACKGROUND: Structured data codes capture acute bodily injury from firearm violence but do not necessarily describe follow-up care from bodily injury and secondary exposure to firearm violence (eg, witnessing a shooting, being threatened by a firearm, or losing a loved one to gun violence and injury from firearms) even though such exposure is associated with many short- and long-term health impacts. Clinical notes from electronic health records (EHRs) often contain data not otherwise captured in structured data fields and can be categorized using natural language processing (NLP). OBJECTIVE: This study protocol outlines the steps being taken to develop an NLP text classifier for determination of exposure to firearm violence (both primary and secondary exposure) from ambulatory primary care and behavioral health EHR clinical notes for persons aged ≥5 years. METHODS: The study will use unstructured data from clinical notes taken between 2012 and 2022 from OCHIN, a multistate network of community health organizations using a single instance of Epic EHR. We describe the process of developing a labeled dataset for supervised NLP development that includes establishing a lexicon (words related to firearm violence) to identify potentially relevant notes, followed by a review of text extracted from a sample of these notes. We then describe the process of building, training, and evaluating candidate machine learning, neural network, and large language model NLP text classifiers. From this, a final NLP model is chosen then evaluated on a new set of randomly selected notes. An engaged stakeholder advisory committee will provide input and guidance on methods and results to identify and address potential biases in the NLP text classifiers. RESULTS: The study was funded in September 2023. Study activities have been ongoing through July 2025 and we are currently evaluating NLP text classifiers. We expect that the final model will be selected by August 2025 and we will publish results of NLP model development and the final model performance in 2026. CONCLUSIONS: This work describes the development of a novel NLP text classifier to identify exposure to firearm violence in ambulatory primary care and behavioral health clinical notes. The NLP model developed in this study may lead to increased ascertainment of patients with exposure, laying the groundwork for understanding the long-term impacts and outcomes of firearm violence exposure and presenting opportunities for improved patient care. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/76681.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".