Under-reporting of pertussis in Ontario: A Canadian Immunization Research Network (CIRN) study using capture-recapture
Bibliographic record
Abstract
INTRODUCTION: Under-reporting of pertussis cases is a longstanding challenge. We estimated the true number of pertussis cases in Ontario using multiple data sources, and evaluated the completeness of each source. METHODS: We linked data from multiple sources for the period 2009 to 2015: public health reportable disease surveillance data, public health laboratory data, and health administrative data (hospitalizations, emergency department visits, and physician office visits). To estimate the total number of pertussis cases in Ontario, we used a three-source capture-recapture analysis stratified by age (infants, or aged one year and older) and adjusting for dependency between sources. We used the Bayesian Information Criterion to compare models. RESULTS: Using probable and confirmed reported cases, laboratory data, and combined hospitalizations/emergency department visits, the estimated total number of cases during the six-year period amongst infants was 924, compared with 545 unique observed cases from all sources. Using the same sources, the estimated total for those aged 1 year and older was 12,883, compared with 3,304 observed cases from all sources. Only 37% of infants and 11% for those aged 1 year and over admitted to hospital or seen in an emergency department for pertussis were reported to public health. Public health reporting sensitivity varied from 2% to 68% depending on age group and the combination of data sources included. Sensitivity of combined hospitalizations and emergency department visits varied from 37% to 49% and of laboratory data from 1% to 50%. CONCLUSIONS: All data sources contribute cases and are complementary, suggesting that the incidence of pertussis is substantially higher than suggested by routine reports. The sensitivity of different data sources varies. Better case identification is required to improve pertussis control in Ontario.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".