Syndrome classification through a retrospective analysis of porcine submissions to a regional animal health laboratory
Bibliographic record
Abstract
In response to the global threats of emerging infectious diseases and bioterrorism events, public health surveillance developed analytical methods to cluster early health indicators from multiple data sources into “syndromes” for rapid and efficient disease detection. Syndromic surveillance has become well established in public health, using many different health indicators from multiple sources. In animal health, the timeliness and efficiency of disease detection in early warning surveillance systems has been enhanced by including syndromic surveillance methods. Animal health syndromic surveillance improves disease detection through the analysis of pre-diagnostic data collected for other purposes, from sources such as laboratories, veterinary clinics, abattoirs, farms and pharmacies. However, the data are inherently non-disease specific compared to traditional surveillance and require analyses to ensure that syndromes represent significant diseases as accurately as possible. Syndrome classification is an analytical process that identifies, collates and validates pre-diagnostic indicators within a data source into accurate and viable syndromes.\nThe goals of this thesis were as follows: a) Review surveillance systems and methods to understand the scale, complexity and validity of different syndromic surveillance approaches. b) Describe and evaluate six years of swine laboratory submission data to Veterinary Diagnostic Services (VDS) in the province of Manitoba, Canada, for the purpose of syndromic surveillance. c) Finally, identify and validate the most appropriate syndromes from pre-diagnostic data within the submitted swine cases.\nAn initial systematic review of public health syndromic surveillance was conducted with 81 studies meeting the criteria. The variety and frequency of populations under surveillance, information sources, pre-diagnostic indicators, syndromes and reported values were recorded. The predominant methods for syndrome classification, temporal and spatial analysis and aberration detection were also described.\n21,665 swine laboratory submissions from January 2003 to March 2009, including 4726 pathology cases, were evaluated. The frequency and distributions of the predominant pre-diagnostic indicators, test requests and specimen types, were described. The most common pathology diagnoses and organ system involvement were reported for the pathology submissions. For syndrome validation, a Multiple Correspondence Analysis was conducted to cluster multiple pathology diagnoses per case into four diagnostic groups based on organ systems; Respiratory, Multisystemic, Gastrointestinal and “Other”.\nSyndrome classification was completed, first using agglomerative hierarchical clustering to classify syndromes from 30 test requests and 34 specimen types. For validation, the syndromes were used as predictive variables in a multinomial logistic regression model applied to training and test data sets. The overall model sensitivity, specificity and predictive values for each organ system outcome were estimated. The individual syndromes were compared using relative risk ratios and marginal effects. Five syndromes were identified as having a significantly higher predictive association with one organ system group (compared to the other three): Respiratory, GI, Reproductive, Joint and PCV (specific to porcine circovirus associated disease).\nThe methods in this thesis identified a simplified analytical approach for syndrome classification of laboratory test requests and specimen types within swine submissions. Alternative algorithms for syndrome grouping, establishment of temporal baselines and exploration of automated aberration detection were identified as areas for future research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.017 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".