Estimating the Epidemic Size of Superspreading Coronavirus Outbreaks in Real Time: Quantitative Study
Bibliographic record
Abstract
BACKGROUND: Novel coronaviruses have emerged and caused major epidemics and pandemics in the past 2 decades, including SARS-CoV-1, MERS-CoV, and SARS-CoV-2, which led to the current COVID-19 pandemic. These coronaviruses are marked by their potential to produce disproportionally large transmission clusters from superspreading events (SSEs). As prompt action is crucial to contain and mitigate SSEs, real-time epidemic size estimation could characterize the transmission heterogeneity and inform timely implementation of control measures. OBJECTIVE: This study aimed to estimate the epidemic size of SSEs to inform effective surveillance and rapid mitigation responses. METHODS: We developed a statistical framework based on back-calculation to estimate the epidemic size of ongoing coronavirus SSEs. We first validated the framework in simulated scenarios with the epidemiological characteristics of SARS, MERS, and COVID-19 SSEs. As case studies, we retrospectively applied the framework to the Amoy Gardens SARS outbreak in Hong Kong in 2003, a series of nosocomial MERS outbreaks in South Korea in 2015, and 2 COVID-19 outbreaks originating from restaurants in Hong Kong in 2020. RESULTS: The accuracy and precision of the estimation of epidemic size of SSEs improved with longer observation time; larger SSE size; and more accurate prior information about the epidemiological characteristics, such as the distribution of the incubation period and the distribution of the onset-to-confirmation delay. By retrospectively applying the framework, we found that the 95% credible interval of the estimates contained the true epidemic size after 37% of cases were reported in the Amoy Garden SARS SSE in Hong Kong, 41% to 62% of cases were observed in the 3 nosocomial MERS SSEs in South Korea, and 76% to 86% of cases were confirmed in the 2 COVID-19 SSEs in Hong Kong. CONCLUSIONS: Our framework can be readily integrated into coronavirus surveillance systems to enhance situation awareness of ongoing SSEs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.053 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".