From Wastewater to Infection Estimates: Incident COVID-19 Infections during Omicron in the U.S.
Bibliographic record
Abstract
Abstract Reconstructing the course of the COVID-19 pandemic through infection estimates is important for assessing disease burden and characterizing transmission dynamics. While wastewater concentration data have been used to estimate infections in localized pre-Omicron studies, a scalable approach that incorporates reinfections and variant-specific shedding remains underdeveloped. To this end, we develop a multi-source approach to retrospectively estimate daily COVID-19 infections in U.S. states during the Omicron era. Our approach integrates wastewater and seroprevalence surveillance data and further incorporates state-specific reinfections to improve infection estimates during the Delta-Omicron transition period. These refined estimates, along with wastewater concentration data adjusted for limited coverage, are used to calculate variant-specific shedding rates, which inform daily infection estimates going forward. While case-based estimates tend to exhibit striking volatility, these infection estimates show more stable and interpretable patterns that closely align with Omicron subvariant transitions. Moreover, we directly quantify the degree of underreporting, showing the extent that reported cases significantly underestimate disease burden, with the lowest reporting rates of 9.72% in Washington, 9.73% in Minnesota, and 10.70% in New York. In the states under study, case reports capture less than a quarter of total infections, leaving the vast majority unaccounted for in official reports. Furthermore, reporting rates differ markedly across states, with disparities growing over time, reflecting the overall rise in underreporting and location-specific limitations to surveillance accuracy. Finally, we estimate time-varying effective reproduction numbers and growth rates to provide a more accurate and timely picture of transmission dynamics over the Omicron era in U.S. states.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".