Drivers of Uncertainty in Estimates of Foodborne Gastroenteritis Incidence
Bibliographic record
Abstract
BACKGROUND: Estimates of the incidence of foodborne illness are increasingly used at national and international levels to quantify the burden of disease and advocate for improvements in food safety. The calculation of such estimates involves multiple datasets and several disease multipliers, applied to dozens of pathogens. Unsurprisingly, this process often produces wide interval estimates. MATERIALS AND METHODS: Using a model of foodborne gastroenteritis in Australia, we calculate the contribution of both data and multipliers to the width of the interval. We then compare pathogen-specific estimates of the proportion of gastroenteritis that is foodborne from national-level studies conducted in Canada, Greece, France, the Netherlands, New Zealand, the United Kingdom, and the United States. RESULTS: Overall, we estimate that 74% (range 63-92%) of the interval width for foodborne gastroenteritis in Australia is a result of uncertainty in the proportion of gastroenteritis that is due to contaminated food. Across national studies, we find considerable variability in point estimates and the width of interval estimates for the foodborne proportion for relatively common pathogens such as Salmonella spp., Campylobacter spp., and norovirus. CONCLUSIONS: While some uncertainty in estimates of gastroenteritis incidence is inevitable, an understanding of the drivers of this uncertainty can help to focus further research. In particular, this work highlights the value of studies quantifying the routes of transmission for common pathogens.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.023 | 0.099 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".