Comparative evaluation of clinical and wastewater genomic surveillance for SARS-CoV-2: implications for integrated infectious disease monitoring
Notice bibliographique
Résumé
ABSTRACT Introduction The COVID-19 pandemic demonstrated the need for comprehensive, cost-effective surveillance systems integrating multiple data streams. This study directly compares SARS-CoV-2 genomic data from wastewater-based surveillance (WBS) and clinical diagnostic testing (CDT) to evaluate lineage diversity, detection timing, and persistence patterns that could inform integrated surveillance strategies. Methods We analyzed SARS-CoV-2 genomic data from Alberta, Canada (July 2022 – March 2025), encompassing 13 municipal wastewater treatment plants covering 80% of the provincial population and clinical samples from provincial diagnostic testing. Clinical samples (n=28,610) and wastewater samples (n=1,685) were sequenced using a tiled amplicon approach. We compared lineage richness over time, lead time for first detection using collection dates and explored four additional wastewater metrics: abundance at first detection, peak abundance, time to peak abundance, and total time detected. The comparison grouped lineages into those found only in WBS and those that were seen in both WBS and CDT. Results Of the 2,586 unique lineages identified over the study period, 1,588 (61.1%) appeared exclusively in WBS, 42 (1.6%) only in CDT, and 956 (36.9%) in both systems. WBS consistently demonstrated higher monthly lineage richness (95-660 lineages) compared to CDT (23-160 lineages). While WBS detected lineages an average of almost 12 days earlier than CDT, the most frequent pattern showed CDT detection first by 8 days, indicating substantial variability. Lineages detected in both systems showed significantly higher initial relative abundance, peak relative abundance, longer persistence, and delayed time to peak compared to WBS-only lineages (all p<0.0001). Conclusions WBS and CDT provide complementary surveillance capabilities with distinct strengths. Rather than relying on “ first detection” as an early warning metric, integrated surveillance should prioritize concordance patterns and abundance metrics that indicate lineages with sustainable transmission potential. These findings support developing surveillance frameworks that strategically combine population-level WBS monitoring with case-linked CDT data for more effective public health response. KEY MESSAGES What is already known on this topic WBS can detect SARS-CoV-2 and other pathogens at the population level and has been described as an early warning system, while CDT provides case-linked genomic data but may be biased towards symptomatic, high-risk or healthcare-seeking individuals. Both surveillance methods were used during the COVID-19 pandemic, but it remains uncertain how to best integrate WBS and CDT genomics and which WBS metrics are most informative. What this study adds This study provides the first direct comparison of SARS-CoV-2 genomic data from WBS and CDT processed by a single laboratory over nearly three years. WBS consistently captured greater lineage diversity than CDT, but “first detection” showed substantial variability, with lineages frequently detected in clinical samples before wastewater. Lineages identified in both surveillance systems showed distinct signatures: higher peak abundance and longer persistence, suggesting these concordance patterns are more reliable indicators than timing alone. How this study might affect research, practice or policy Integrated surveillance frameworks should prioritize monitoring concordance between WBS and CDT, using abundance and persistence metrics to identify lineages with significant transmission potential. The complementary strengths of these systems support their combined use for cost-effective surveillance with WBS monitoring population trends and CDT providing clinical context for targeted interventions.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,003 | 0,006 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».