A metagenomics workflow for SARS-CoV-2 identification, co-pathogen detection, and overall diversity
Bibliographic record
Abstract
An unbiased metagenomics approach to virus identification can be essential in the initial phase of a pandemic. Better molecular surveillance strategies are needed for the detection of SARS-CoV-2 variants of concern and potential co-pathogens triggering respiratory symptoms. Here, a metagenomics workflow was developed to identify the metagenome diversity by SARS-CoV-2 diagnosis (npositive = 65; nnegative = 60), symptomatology status (nsymptomatic = 71; nasymptomatic = 54) and anatomical swabbing site (nnasopharyngeal = 96; nthroat = 29) in 125 individuals. Furthermore, the workflow was able to identify putative respiratory co-pathogens, and the SARS-CoV-2 lineage across 29 samples. The diversity analysis showed a significant shift in the DNA-metagenome by symptomatology status and anatomical swabbing site. Additionally, metagenomic diversity differed between SARS-CoV-2 infected and uninfected asymptomatic individuals. While 31 co-pathogens were identified in SARS-CoV-2 infected patients, no significant increase in pathogen or associated reads were noted when compared to SARS-CoV-2 negative patients. The Alpha SARS-CoV-2 VOC and 2 variants of interest (Zeta) were successfully identified for the first time using a clinical metagenomics approach. The metagenomics pipeline showed a sensitivity of 86% and a specificity of 72% for the detection of SARS-CoV-2. Clinical metagenomics can be employed to identify SARS-CoV-2 variants and respiratory co-pathogens potentially contributing to COVID-19 symptoms. The overall diversity analysis suggests a complex set of microorganisms with different genomic abundance profiles in SARS-CoV-2 infected patients compared to healthy controls. More studies are needed to correlate severity of COVID-19 disease in relation to potential disbyosis in the upper respiratory tract. A metagenomics approach is particularly useful when novel pandemic pathogens emerge.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".