MétaCan
Menu
Back to cohort
Record W7110074332 · doi:10.64898/2025.12.02.25341504

Comparative evaluation of clinical and wastewater genomic surveillance for SARS-CoV-2: implications for integrated infectious disease monitoring

2025· article· W7110074332 on OpenAlexafffundabout

Bibliographic record

VenuemedRxiv · 2025
Typearticle
Language
FieldMedicine
TopicSARS-CoV-2 detection and testing
Canadian institutionsPublic Health Agency of CanadaUniversity of AlbertaUniversity of CalgaryAlberta Health Services
FundersPublic Health AgencyPublic Health Agency of Canada
KeywordsLineage (genetic)Species richnessInfectious disease (medical specialty)WastewaterPopulationPandemicGenomics

Abstract

fetched live from OpenAlex

ABSTRACT Introduction The COVID-19 pandemic demonstrated the need for comprehensive, cost-effective surveillance systems integrating multiple data streams. This study directly compares SARS-CoV-2 genomic data from wastewater-based surveillance (WBS) and clinical diagnostic testing (CDT) to evaluate lineage diversity, detection timing, and persistence patterns that could inform integrated surveillance strategies. Methods We analyzed SARS-CoV-2 genomic data from Alberta, Canada (July 2022 – March 2025), encompassing 13 municipal wastewater treatment plants covering 80% of the provincial population and clinical samples from provincial diagnostic testing. Clinical samples (n=28,610) and wastewater samples (n=1,685) were sequenced using a tiled amplicon approach. We compared lineage richness over time, lead time for first detection using collection dates and explored four additional wastewater metrics: abundance at first detection, peak abundance, time to peak abundance, and total time detected. The comparison grouped lineages into those found only in WBS and those that were seen in both WBS and CDT. Results Of the 2,586 unique lineages identified over the study period, 1,588 (61.1%) appeared exclusively in WBS, 42 (1.6%) only in CDT, and 956 (36.9%) in both systems. WBS consistently demonstrated higher monthly lineage richness (95-660 lineages) compared to CDT (23-160 lineages). While WBS detected lineages an average of almost 12 days earlier than CDT, the most frequent pattern showed CDT detection first by 8 days, indicating substantial variability. Lineages detected in both systems showed significantly higher initial relative abundance, peak relative abundance, longer persistence, and delayed time to peak compared to WBS-only lineages (all p<0.0001). Conclusions WBS and CDT provide complementary surveillance capabilities with distinct strengths. Rather than relying on “ first detection” as an early warning metric, integrated surveillance should prioritize concordance patterns and abundance metrics that indicate lineages with sustainable transmission potential. These findings support developing surveillance frameworks that strategically combine population-level WBS monitoring with case-linked CDT data for more effective public health response. KEY MESSAGES What is already known on this topic WBS can detect SARS-CoV-2 and other pathogens at the population level and has been described as an early warning system, while CDT provides case-linked genomic data but may be biased towards symptomatic, high-risk or healthcare-seeking individuals. Both surveillance methods were used during the COVID-19 pandemic, but it remains uncertain how to best integrate WBS and CDT genomics and which WBS metrics are most informative. What this study adds This study provides the first direct comparison of SARS-CoV-2 genomic data from WBS and CDT processed by a single laboratory over nearly three years. WBS consistently captured greater lineage diversity than CDT, but “first detection” showed substantial variability, with lineages frequently detected in clinical samples before wastewater. Lineages identified in both surveillance systems showed distinct signatures: higher peak abundance and longer persistence, suggesting these concordance patterns are more reliable indicators than timing alone. How this study might affect research, practice or policy Integrated surveillance frameworks should prioritize monitoring concordance between WBS and CDT, using abundance and persistence metrics to identify lineages with significant transmission potential. The complementary strengths of these systems support their combined use for cost-effective surveillance with WBS monitoring population trends and CDT providing clinical context for targeted interventions.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.006
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.047
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.006
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.186
GPT teacher head0.459
Teacher spread0.273 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes3
Has abstractyes

Explore more

Same venuemedRxivSame topicSARS-CoV-2 detection and testingFrench-language works237,207