Development and validation of an automated emergency department-based syndromic surveillance system to enhance public health surveillance in Yukon: a lower-resourced and remote setting
Bibliographic record
Abstract
BACKGROUND: Automated Emergency Department syndromic surveillance systems (ED-SyS) are useful tools in routine surveillance activities and during mass gathering events to rapidly detect public health threats. To improve the existing surveillance infrastructure in a lower-resourced rural/remote setting and enhance monitoring during an upcoming mass gathering event, an automated low-cost and low-resources ED-SyS was developed and validated in Yukon, Canada. METHODS: Syndromes of interest were identified in consultation with the local public health authorities. For each syndrome, case definitions were developed using published resources and expert elicitation. Natural language processing algorithms were then written using Stata LP 15.1 (Texas, USA) to detect syndromic cases from three different fields (e.g., triage notes; chief complaint; discharge diagnosis), comprising of free-text and standardized codes. Validation was conducted using data from 19,082 visits between October 1, 2018 to April 30, 2019. The National Ambulatory Care Reporting System (NACRS) records were used as a reference for the inclusion of International Classification of Disease, 10th edition (ICD-10) diagnosis codes. The automatic identification of cases was then manually validated by two raters and results were used to calculate positive predicted values for each syndrome and identify improvements to the detection algorithms. RESULTS: A daily secure file transfer of Yukon's Meditech ED-Tracker system data and an aberration detection plan was set up. A total of six syndromes were originally identified for the syndromic surveillance system (e.g., Gastrointestinal, Influenza-like-Illness, Mumps, Neurological Infections, Rash, Respiratory), with an additional syndrome added to assist in detecting potential cases of COVID-19. The positive predictive value for the automated detection of each syndrome ranged from 48.8-89.5% to 62.5-94.1% after implementing improvements identified during validation. As expected, no records were flagged for COVID-19 from our validation dataset. CONCLUSIONS: The development and validation of automated ED-SyS in lower-resourced settings can be achieved without sophisticated platforms, intensive resources, time or costs. Validation is an important step for measuring the accuracy of syndromic surveillance, and ensuring it performs adequately in a local context. The use of three different fields and integration of both free-text and structured fields improved case detection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".