Updating global estimates of pathogen-attributable diarrhoeal disease burden: a methodology and integrated protocol for a broad-scope systematic review of a syndrome with diverse infectious aetiologies
Bibliographic record
Abstract
Introduction Sustaining declines in global infectious disease burden will increasingly require efforts targeted to specific aetiological agents and common transmission pathways, particularly in this era of global change and human interconnectivity accelerating transmission and emergence of infectious pathogens. Systematic reviews and meta-analyses can be an effective and resource-efficient method for synthesising evidence regarding disease epidemiology for a wide range of pathogens and are the evidence source used by initiatives like the Planetary Child Health and Enterics Observatory (Plan-EO) and the WHO to determine the aetiology-specific epidemiology of diarrhoeal disease. Therefore, we developed this integrated systematic review methodology and protocol that aims to compile a database of published prevalence estimates for 17 diarrhoea-causing pathogens as inputs for disease burden estimation. Methods and analysis We will seek estimates of the prevalence of each endemic enteric pathogen estimated from published population-based studies that diagnosed their presence in stool samples from both asymptomatic subjects and those experiencing diarrhoea. The pathogens include the enteric viruses adenovirus, astrovirus, norovirus, rotavirus and sapovirus, the bacteria Campylobacter , Shigella , Salmonella enterica , Vibrio cholerae and the Escherichia coli (E. coli) pathotypes enteroaggregative E. coli , enteropathogenic E. coli , enterotoxigenic E. coli and Shiga-toxin-producing E. coli and the intestinal protozoa Cryptosporidium , Cyclospora , Entamoeba histolytica and Giardia . Meta-analytical methods for analyses of the resulting database (including risk of bias analysis) will be published alongside their findings. Ethics and dissemination This systematic review is exempt from ethics approval because the work is carried out on published documents. The database that results from this review will be made available as a supplementary file of the resulting published manuscript. It will also be made available for download from the Plan-EO website, where updated versions will be posted on a quarterly basis. PROSPERO registration number CRD42023427998.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.315 | 0.449 |
| Meta-epidemiology (narrow) | 0.008 | 0.008 |
| Meta-epidemiology (broad) | 0.021 | 0.025 |
| Bibliometrics | 0.038 | 0.032 |
| Science and technology studies | 0.004 | 0.008 |
| Scholarly communication | 0.010 | 0.009 |
| Open science | 0.008 | 0.013 |
| Research integrity | 0.011 | 0.008 |
| Insufficient payload (model declined to judge) | 0.023 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".