Bibliographic record
Abstract
Abstract Understanding the dynamics of the COVID-19 pandemic, evaluating the efficacy of past and current control measures, and estimating vaccination needs, requires knowledge of the number of infections in the population over time. This number, however, generally differs substantially from the number of confirmed cases due to a large fraction of asymptomatic infections as well as geographically and temporally variable testing effort and strategies. Here I use age-stratified death count statistics, age-dependent infection fatality risks and stochastic modeling to estimate the prevalence and growth of SARS-CoV-2 infections among adults (age ≥ 20 years) in 171 countries, from early 2020 until April 9, 2021. The accuracy of the approach is confirmed through comparison to previous nationwide general-population seroprevalence surveys in multiple countries. Estimates of infections over time, compared to reported cases, reveal that the fraction of infections that are detected vary widely over time and between countries, and hence comparisons of confirmed cases alone (between countries or time points) often yield a false picture of the pandemic’s dynamics. As of April 9, 2021, the nationwide cumulative SARS-CoV-2 prevalence (past and current infections relative to the population size) is estimated at 61% (95%-CI 42-78) for Peru, 58% (39–83) for Mexico, 57% (31–75) for Brazil, 55% (34–72) for South Africa, 29% (19-48) for the US, 26% (16–49) for the United Kingdom, 19% (12–34) for France, 19% (11–33) for Sweden, 9.6% (6.5–15) for Canada, 11% (7–19) for Germany and 0.67% (0.47–1.1) for Japan. The presented time-resolved estimates expand the possibilities to study the factors that influenced and still influence the pandemic’s progression in 171 countries. Regular updates are available at: www.loucalab.com/archive/COVID19prevalence
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".