MétaCan
Menu
Retour à la cohorte
Enregistrement W6930636582 · doi:10.5281/zenodo.15183896

Supplementary Data - Comparative metagenomics of two shallow marine microbial communities in western Greenland - Raw Sequence Data

2023· dataset· en· W6930636582 sur OpenAlexaffabout

Notice bibliographique

RevueZenodo (CERN European Organization for Nuclear Research) · 2023
Typedataset
Langueen
DomaineNursing
ThématiqueMagnesium in Health and Disease
Établissements canadiensMcMaster University
Organismes subventionnairesnon disponible
Mots-clésMetagenomicsEnvironmental DNASeawaterSampling (signal processing)FjordBiomass (ecology)Raw data

Résumé

récupéré en direct d'OpenAlex

Raw (cleaned) reads associated with manuscript "Comparative metagenomics of two shallow marine microbial communities in western Greenland". Samples were collected under a non-exclusive licence for non-commercial utilization of Greenland genetic resources (licence no. G23-045), provided by the Government of Greenland to Daniel G. Dick. See attached "TERMS AND CONDITIONS FOR LICENCES FOR UTILIZATION OF GREENLAND GENETIC RESOURCES" for a complete description of the terms associated with the use of this data. Methods: Free-floating microorganisms and environmental DNA (eDNA) were collected from two localities in western Greenland (Sisimiut [66.9369, -53.6981] and Ilulissat [69.2200, -51.1121]) between August 16th and August 17th, 2023, using a Waterra eDNA Sampling Pump (Figure 1). Samples were collected from shore, with seawater being collected at the surface (<10 cm). Sampling was conducted at a third locality (Kangerlussuaq, 66.9663, -50.9489), however, the resulting samples had significantly fewer reads than the other two localities, potentially due to low biomass in the region and/or the presence of silt in the water which clogged the filters (see Supplementary Information). Due to the low-quality samples from this site, samples from Kangerlussuaq were excluded from this analysis, but are included in the Supplementary Information for reference. For each sample, three litres of seawater was pumped through a sterile 0.22 μm Waterra eDNA filter (five samples were collected from each locality). Following this, 50 ml of Zymo Research DNA/RNA Shield was injected into the filter capsule. Zymo Research DNA/RNA Shield inactivates the collected microorganisms while ensuring soft tissues and eDNA are preserved during transport (Trivedi et al. 2022). In order to dislodge the collected microorganisms and eDNA, the filter capsule was agitated for 60 seconds (with the capsule being rotated once after 30 seconds) using a Waterra eDNA Filter Shaker. Once agitation was complete, the Zymo Research DNA/RNA Shield containing preserved microorganisms and eDNA was transferred to a sterile 50 ml Falcon test tube. The Falcon tubes containing collected microorganisms and eDNA preserved in DNA/RNA Shield were then frozen at -20°C for the remainder of the voyage (use of DNA/RNA Shield precluded the need to store samples at colder temperatures – see Trivedi et al. [2022] for a review of the use of DNA/RNA Shield with samples from the polar regions). Samples were collected under a non-exclusive licence for non-commercial utilization of Greenland genetic resources (licence no. G23-045), provided by the Government of Greenland. Quantitative PCR and shotgun metagenomic analyses were conducted at Microbiome Insights (Richmond, British Columbia, Canada) using the following protocol. DNA was extracted using the Qiagen MagAttract PowerSoil DNA KF kit (Formerly MOBio PowerSoil DNA Kit) using a KingFisher robot. DNA quality was evaluated visually via gel electrophoresis and quantified using a Qubit 3.0 fluorometer (Thermo-Fischer, Waltham, MA, USA). Libraries were prepared using an Illumina Nextera library preparation kit with an in-house protocol (Illumina, San Diego, CA, USA). Paired-end sequencing (150 bp x 2) was done in a NovaSeq 6000 instrument. Shotgun metagenomic sequence reads were processed with the Sunbeam pipeline (Clarke et al. 2019). Initial quality evaluation was done using FastQC v0.11.5 (Clarke et al. 2019). Processing took part in four steps: adapter removal, read trimming, low-complexity-reads removal, and host-sequence removals. Adapter removal was done using cutadapt v2.6 (Martin 2015). Trimming was done with Trimmomatic v0.36 using custom parameters (LEADING:3 TRAILING:3 SLIDINGWINDOW:4:15 MINLEN:36) (Bolger et al. 2014). Low-complexity sequences were detected with Komplexity v0.3.6 (Clarke et al. 2019). High-quality reads were mapped to the human genome (Telomere-to-Telomere assembly, T2T-CHM13v2.0, GCF_009914755.1) and those that mapped to it were removed from the analysis. The remaining reads were taxonomically classified using Kraken2 with the PlusPF database version 2022-09-26 (Wood et al. 2019). For functional profiling, high-quality (filtered) reads were aligned against the SEED database (Overbeek et al. 2014) via translated homology search and annotated to Subsystems, or functional levels, 1-3 using Super-Focus (Silva et al. 2016). Functional gene abundance was expressed using pseudocounts (wherein counts of an individual functional gene are divided by the number of hits in the complete database) following the method described in Silva et al. (2016). In order to account for differences in sampling depth and composition between different samples, DESeq2 normalization (geometric means – Love et al. 2014) was used. To quantify differences in the distribution of taxa and functional genes between localities, pairwise dissimilarities were calculated using the Bray-Curtis index (Bray and Curtis 1957), and samples were compared using permutational multivariate analysis of variance (PERMANOVA – Anderson 2001), using functions from the vegan R package (Oksanen et al. 2024). Comparison of functional genes were made using the Level 3 SEED Subsystem (Overbeek et al. 2014). Alpha diversity of samples was quantified using the Shannon index (Shannon 1948) and compared between locations using the Kruskal-Wallis chi-squared test (Kruskal and Wallis 1952). Finally, to better understand the differences between these two localities, Wald tests (Wald 1943) were used to test for significant differences in the abundance of specific taxa and the expression of specific functional genes (using the DESeq2 R package – Love et al. 2014).

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,007
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesCharge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Jeu de données · Signal consensuel: Jeu de données
Score de désaccord entre enseignants0,349
Score d'incertitude au seuil0,928

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0020,007
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0040,007
Études des sciences et des technologies0,0020,000
Communication savante0,0020,001
Science ouverte0,0020,001
Intégrité de la recherche0,0010,002
Charge utile insuffisante (le modèle a refusé de juger)0,3490,091

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,239
Tête enseignante GPT0,377
Écart entre enseignants0,137 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeSans objet
Domainenon disponible
GenreJeu de données

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2023
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revueZenodo (CERN European Organization for Nuclear Research)Même sujetMagnesium in Health and DiseaseTravaux en français237 207