MétaCan
Menu
← Back to cohort
Record W4224247167 · doi:10.1101/2022.04.19.488797

High resolution shotgun metagenomics: the more data, the better?

2022· preprint· en· W4224247167 on OpenAlexafffund
Julien Tremblay, Lars Schreiber, Charles W. Greer

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2022
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenomics and Phylogenetic Studies
Canadian institutionsNational Research Council Canada
FundersGraduate and Professional Student Council, University of ArizonaCompute Canada
KeywordsMetagenomicsShotgun sequencingWorkflowDeep sequencingDNA sequencingSample (material)Computer scienceShotgunData miningComputational biologyBiologyGenomeDatabaseGeneticsGene

Abstract

fetched live from OpenAlex

Abstract In shotgun metagenomics (SM), the state of the art bioinformatic workflows are referred to as high resolution shotgun metagenomics (HRSM) and require intensive computing and disk storage resources. While the increase in data output of the latest iteration of high throughput DNA sequencing systems can allow for unprecedented sequencing depth at a minimal cost, adjustments in HRSM workflows will be needed to properly process these ever-increasing sequence datasets. One potential adaptation is to generate so-called shallow SM datasets that contain fewer sequencing data per sample as compared to the more classic high coverage sequencing. While shallow sequencing is a promising avenue for SM data analysis, detailed benchmarks using real data are lacking. In this case study, we took four public SM datasets, one massive and the others moderate in size and subsampled each dataset at various levels to mimic shallow sequencing datasets of various sequencing depths. Our results suggest that shallow SM sequencing is a viable avenue to obtain sound results regarding microbial community structures and that high depth sequencing does not bring additional elements for ecological interpretation. More specifically, results obtained by subsampling as little as 0.5M sequencing clusters per sample were similar to the results obtained with the largest subsampled dataset for the human gut and agricultural soil datasets. For the Antarctic dataset, which contained only a few samples, 4M sequencing clusters per sample was found to generate comparable results to the full dataset. One area where ultra-deep sequencing and maximizing the usage of all data was undeniably beneficial was in the generation of metagenome-assembled genomes (MAGs). Key points – Three public multi-sample shotgun metagenomic NovaSeq datasets totalling 12,389,583 and 202 Gb, respectively were analyzed at various sequencing depths to evaluate the accuracy of shallow shotgun metagenomic sequencing using a high resolution shotgun metagenomic bioinformatic workflow. A synthetic mock community of 20 bacterial genomes was also analyzed for validation purposes. – Datasets subsampled to low sequencing depths gave nearly identical ecological patterns (taxonomic and functional composition and beta-alpha-diversity) compared to high depth subsampled datasets. – Rare taxa and functions could be uncovered with high sequencing depth vs. low sequencing depth datasets, but did not affect global ecological patterns. – High sequencing depth was positively correlated with both quantity and quality of recovered metagenome-assembled genomes.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.008
metaresearch head score (Gemma)0.009
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.008
Threshold uncertainty score0.045

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0080.009
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.003
Science and technology studies0.0010.001
Scholarly communication0.0040.004
Open science0.0010.002
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.021
GPT teacher head0.231
Teacher spread0.210 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2022
Admission routes2
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)→Same topicGenomics and Phylogenetic Studies→French-language works237,207→