All together now: aggregating multiple records to develop a person-based dataset to integrate and enhance infectious disease surveillance in Ontario, Canada
Bibliographic record
Abstract
SETTING: Syndemics occur when two or more health conditions interact to increase morbidity and mortality and are exacerbated by social, economic, environmental, and political factors. Routine provincial surveillance in Ontario assesses and reports on the epidemiology of single infectious diseases separately. Therefore, we aimed to develop a method that allows disease overlaps to be examined routinely as a path to better understanding and addressing syndemics in Ontario. INTERVENTION: We extracted data for individuals with a record of chlamydia, gonorrhea, infectious syphilis, hepatitis B and C, HIV/AIDS, invasive group A streptococcal disease (iGAS), or tuberculosis in Ontario's reportable disease database from 1990 to 2018. We transformed the data into a person-based integrated surveillance dataset retaining individuals (clients) with at least one record between 2006 and 2018. OUTCOMES: The resulting dataset had 659,136 unique disease records among 470,673 unique clients. Of those clients, 23.1% had multiple disease records with 50 being the most for one client. We described the frequency of disease overlaps; for example, 34.7% of clients with a syphilis record had a gonorrhea record. We quantified known overlaps, finding 1274 clients had gonorrhea, infectious syphilis, and HIV/AIDS records, and potentially emerging overlaps, finding 59 clients had HIV/AIDS, hepatitis C, and iGAS records. IMPLICATIONS: Our novel person-based integrated surveillance dataset represents a platform for ongoing in-depth assessment of disease overlaps such as the relative timing of disease records. It enables a more client-focused approach, is a step towards improved characterization of syndemics in Ontario, and could inform other jurisdictions interested in adopting similar approaches.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".