Timeliness of reporting of SARS-CoV-2 seroprevalence results and their utility for infectious disease surveillance
Bibliographic record
Abstract
Seroprevalence studies have been used throughout the COVID-19 pandemic to monitor infection and immunity. These studies are often reported in peer-reviewed journals, but the academic writing and publishing process can delay reporting and thereby public health action. Seroprevalence estimates have been reported faster in preprints and media, but with concerns about data quality. We aimed to (i) describe the timeliness of SARS-CoV-2 serosurveillance reporting by publication venue and study characteristics and (ii) identify relationships between timeliness, data validity, and representativeness to guide recommendations for serosurveillance efforts. We included seroprevalence studies published between January 1, 2020 and December 31, 2021 from the ongoing SeroTracker living systematic review. For each study, we calculated timeliness as the time elapsed between the end of sampling and the first public report. We evaluated data validity based on serological test performance and correction for sampling error, and representativeness based on the use of a representative sample frame and adequate sample coverage. We examined how timeliness varied with study characteristics, representativeness, and data validity using univariate and multivariate Cox regression. We analyzed 1844 studies. Median time to publication was 154 days (IQR 64-255), varying by publication venue (journal articles: 212 days, preprints: 101 days, institutional reports: 18 days, and media: 12 days). Multivariate analysis confirmed the relationship between timeliness and publication venue and showed that general population studies were published faster than special population or health care worker studies; there was no relationship between timeliness and study geographic scope, geographic region, representativeness, or serological test performance. Seroprevalence studies in peer-reviewed articles and preprints are published slowly, highlighting the limitations of using the academic literature to report seroprevalence during a health crisis. More timely reporting of seroprevalence estimates can improve their usefulness for surveillance, enabling more effective responses during health emergencies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.275 | 0.617 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.008 |
| Bibliometrics | 0.021 | 0.031 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.006 | 0.008 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".