Characteristics of outbreak-response databases: Canadian SARS example
Bibliographic record
Abstract
In 2003, SARS was a serious health concern in Canada. As of September 3 of that year, the Public Health Agency of Canada reported a total of 438 cases: 251 Probable (247 Ontario, 4 British Columbia) and 187 Suspect (128 Ontario, 46 British Columbia). Although the outbreak was short-lived, more than forty people died from the disease. A substantial database evolved as a consequence of control efforts. Specimens began to be received and tested at the National Microbiology Laboratory (NML) in Winnipeg, Manitoba, on March 17, 2003. NML’s SARS database contains more than 12,000 records and 192 variables, with variables detailing clinical/ diagnostic (17), microbiological (143), epidemiological (25), and administrative (7) features. Clinical variables include: diarrhea, difficulty of breathing, severity of illness, systemic status, date of onset of illness, case status, case status modification, and case status modification date. Diagnostic variables include: fever, chest X-ray change, cough, shortness of breath, contact with probable case, travel, source of exposure, and contact type. Epidemiological data include: date of birth, age, sex, epidemiology cluster, and employment status. Administrative data include: patient’s last name, first name, temporal data (date of collection of the specimen, date of its receipt, and first date of hospitalization of the patient) and spatial data (origin of specimen, and identity of the hospital). A wide variety of laboratory tests are included in the database: Enzyme-Linked Immunosorbent Assay (ELISA, 7 variables), Immunofluorescence Assay (IFA, 7); plaque reduction neutralization test (PRN, 3); cytopathogenic test (CPE, 2), and electron microscopy test (EM, 8). There are tests for Coronavirus (13 variables), human metapneumovirus (hMPV, 16), circovirus (Circo, 4), porcine circovirus (PCV1, 6), TTvirus (TTV, 7), TTV-like-mini-virus (TLMV, 7), Hantaanvirus (1), Rhinovirus (3), and Paramyxovirus (3). Nested PCR, RT-PCR, and sequencing tests are common among the viruses. The NML-SARS database evolved as part of an ongoing effort involving multiple institutions, multiple regions, intense time pressures, and the participation of many operational and scientific specialties (e.g., clinicians, epidemiologists, microbiologists, administrators). Although the database arose in response to a specific disease in Canada, it can be looked upon as an example of what might typically arise from a publichealth response to an outbreak of an emerging disease. Hence there is value in analysing the NML-SARS database, looking for general characteristics, and highlighting where opportunities for scientific advances exist. This is our objective. Putative characteristics of outbreak-response data sets are: ad hoc by definition; evolving database and data administration; basic assumptions (e.g., case definition) open to refinement; and insights that are often merely suggestive. Opportunities for scientific advancement include: Exploratory Data Analysis of evolving data sets; definition of a relational data model appropriate to outbreak-response data sets; improvement of statistical data modelling methodology to estimate empty blocks of cells (resulting as emerging understanding directs interest from one area to another); data analysis to refine basic assumptions made during control operations; process modelling to explore consequences of unverified insights.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.029 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.008 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".