SeroTracker: synthesising seroprevalence data to map the true extent of COVID-19 infection and immunity
Bibliographic record
Abstract
Effective decision-making during pandemics requires understanding the true extent of infection and immunity. Seroprevalence studies measure the population prevalence of antibodies, providing key insight into these quantities. However, the results from these studies are dispersed, hindering their use. We developed SeroTracker, a dashboard and data platform for SARS-CoV-2 seroprevalence studies. Through its ongoing systematic review, SeroTracker identified over 3,800 seroprevalence studies and made their data available to global health, government, and academic stakeholders. Meta-analysing these studies, we show that global SARS-CoV-2 seroprevalence was 59.2% in September 2021 (95% confidence interval 56.1 − 62.2%), and rose steeply during 2021 due to infection in some regions and vaccination in others. We also identify limitations of seroprevalence studies: their timeliness, where the median study is published 154 days after being conducted; measurement bias, where assay manufacturers systematically overestimate sensitivity by 5.4% and specificity by 2.8%; and coverage, where seroprevalence studies are sparse in time and space. To address these challenges, we develop a Bayesian modelling framework that synthesises information from case and seroprevalence data to provide up- to-date estimates of true infections. We validate this model on simulated data and apply it to Canadian data, showing that only approximately 20% of infections were detected as cases at the start of the pandemic, and revealing substantial geographical heterogeneity in infection dynamics. Returning to SeroTracker’s evidence synthesis process, we develop an automated tool for risk of bias assessment with excellent reliability compared to manual review (intraclass correlation 0.77), and implement a natural language processing-based process to screen abstracts for our systematic review that reduces screening time by approximately 70%. Finally, we reflect on the increased risk posed by pandemics and discuss future work applying engineering to public health to reduce these risks, including the development of a central data repository for respiratory virus surveillance data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.061 | 0.238 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.011 |
| Bibliometrics | 0.022 | 0.013 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.013 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".