Estimating internationally imported cases during the early COVID-19 pandemic
Bibliographic record
Abstract
Early in the COVID-19 pandemic, when cases were predominantly reported in the city of Wuhan, China, local outbreaks in Europe, North America, and Asia were largely predicted from imported cases on flights from Wuhan, potentially missing imports from other key source cities. Here, we account for importations from Wuhan and from other cities in China, combining COVID-19 prevalence estimates in 18 Chinese cities with estimates of flight passenger volume to predict for each day between early December 2019 to late February 2020 the number of cases exported from China. We predict that the main source of global case importation in early January was Wuhan, but due to the Wuhan lockdown and the rapid spread of the virus, the main source of case importation from mid February became Chinese cities outside of Wuhan. For destinations in Africa in particular, non-Wuhan cities were an important source of case imports (1 case from those cities for each case from Wuhan, range of model scenarios: 0.1-9.8). Our model predicts that 18.4 (8.5 - 100) COVID-19 cases were imported to 26 destination countries in Africa, with most of them (90%) predicted to have arrived between 7th January (±10 days) and 5th February (±3 days), and all of them predicted prior to the first case detections. We finally observed marked heterogeneities in expected imported cases across those locations. Our estimates shed light on shifting sources and local risks of case importation which can help focus surveillance efforts and guide public health policy during the final stages of the pandemic. We further provide a time window for the seeding of local epidemics in African locations, a key parameter for estimating expected outbreak size and burden on local health care systems and societies, that has yet to be defined in these locations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".