Surveillance bias in the assessment of the size of COVID-19 epidemic waves: a case study
Bibliographic record
Abstract
OBJECTIVES: To estimate the size of COVID-19 waves using four indicators across three pandemic periods and assess potential surveillance bias. STUDY DESIGN: Case study using data from one region of Switzerland. METHODS: We compared cases, hospitalizations, deaths, and seroprevalence during three periods including the first three pandemic waves (period 1: Feb-Oct 2020; period 2: Oct 2020-Feb 2021; period 3: Feb-Aug 2021). Data were retrieved from the Federal Office of Public Health or estimated from population-based studies. To assess potential surveillance bias, indicators were compared to a reference indicator, i.e. seroprevalence during periods 1 and 2 and hospitalizations during the period 3. Timeliness of indicators (the duration from data generation to the availability of the information to decision-makers) was also evaluated. RESULTS: Using seroprevalence (our reference indicator for period 1 and 2), the 2nd wave size was slightly larger (by a ratio of 1.4) than the 1st wave. Compared to seroprevalence, cases largely overestimated the 2nd wave size (2nd vs 1st wave ratio: 6.5), while hospitalizations (ratio: 2.2) and deaths (ratio: 2.9) were more suitable to compare the size of these waves. Using hospitalizations as a reference, the 3rd wave size was slightly smaller (by a ratio of 0.7) than the 2nd wave. Cases or deaths slightly underestimated the 3rd wave size (3rd vs 2nd wave ratio for cases: 0.5; for deaths: 0.4). The seroprevalence was not useful to compare the size of these waves due to high vaccination rates. Across all waves, timeliness for cases and hospitalizations was better than for deaths or seroprevalence. CONCLUSIONS: The usefulness of indicators for assessing the size of pandemic waves depends on the type of indicator and the period of the pandemic.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.032 | 0.070 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".