Supplementary Data - Comparative metagenomics of two shallow marine microbial communities in western Greenland - Raw Sequence Data
Notice bibliographique
Résumé
Raw (cleaned) reads associated with manuscript "Comparative metagenomics of two shallow marine microbial communities in western Greenland". Samples were collected under a non-exclusive licence for non-commercial utilization of Greenland genetic resources (licence no. G23-045), provided by the Government of Greenland to Daniel G. Dick. See attached "TERMS AND CONDITIONS FOR LICENCES FOR UTILIZATION OF GREENLAND GENETIC RESOURCES" for a complete description of the terms associated with the use of this data. Methods: Free-floating microorganisms and environmental DNA (eDNA) were collected from two localities in western Greenland (Sisimiut [66.9369, -53.6981] and Ilulissat [69.2200, -51.1121]) between August 16th and August 17th, 2023, using a Waterra eDNA Sampling Pump (Figure 1). Samples were collected from shore, with seawater being collected at the surface (<10 cm). Sampling was conducted at a third locality (Kangerlussuaq, 66.9663, -50.9489), however, the resulting samples had significantly fewer reads than the other two localities, potentially due to low biomass in the region and/or the presence of silt in the water which clogged the filters (see Supplementary Information). Due to the low-quality samples from this site, samples from Kangerlussuaq were excluded from this analysis, but are included in the Supplementary Information for reference. For each sample, three litres of seawater was pumped through a sterile 0.22 μm Waterra eDNA filter (five samples were collected from each locality). Following this, 50 ml of Zymo Research DNA/RNA Shield was injected into the filter capsule. Zymo Research DNA/RNA Shield inactivates the collected microorganisms while ensuring soft tissues and eDNA are preserved during transport (Trivedi et al. 2022). In order to dislodge the collected microorganisms and eDNA, the filter capsule was agitated for 60 seconds (with the capsule being rotated once after 30 seconds) using a Waterra eDNA Filter Shaker. Once agitation was complete, the Zymo Research DNA/RNA Shield containing preserved microorganisms and eDNA was transferred to a sterile 50 ml Falcon test tube. The Falcon tubes containing collected microorganisms and eDNA preserved in DNA/RNA Shield were then frozen at -20°C for the remainder of the voyage (use of DNA/RNA Shield precluded the need to store samples at colder temperatures – see Trivedi et al. [2022] for a review of the use of DNA/RNA Shield with samples from the polar regions). Samples were collected under a non-exclusive licence for non-commercial utilization of Greenland genetic resources (licence no. G23-045), provided by the Government of Greenland. Quantitative PCR and shotgun metagenomic analyses were conducted at Microbiome Insights (Richmond, British Columbia, Canada) using the following protocol. DNA was extracted using the Qiagen MagAttract PowerSoil DNA KF kit (Formerly MOBio PowerSoil DNA Kit) using a KingFisher robot. DNA quality was evaluated visually via gel electrophoresis and quantified using a Qubit 3.0 fluorometer (Thermo-Fischer, Waltham, MA, USA). Libraries were prepared using an Illumina Nextera library preparation kit with an in-house protocol (Illumina, San Diego, CA, USA). Paired-end sequencing (150 bp x 2) was done in a NovaSeq 6000 instrument. Shotgun metagenomic sequence reads were processed with the Sunbeam pipeline (Clarke et al. 2019). Initial quality evaluation was done using FastQC v0.11.5 (Clarke et al. 2019). Processing took part in four steps: adapter removal, read trimming, low-complexity-reads removal, and host-sequence removals. Adapter removal was done using cutadapt v2.6 (Martin 2015). Trimming was done with Trimmomatic v0.36 using custom parameters (LEADING:3 TRAILING:3 SLIDINGWINDOW:4:15 MINLEN:36) (Bolger et al. 2014). Low-complexity sequences were detected with Komplexity v0.3.6 (Clarke et al. 2019). High-quality reads were mapped to the human genome (Telomere-to-Telomere assembly, T2T-CHM13v2.0, GCF_009914755.1) and those that mapped to it were removed from the analysis. The remaining reads were taxonomically classified using Kraken2 with the PlusPF database version 2022-09-26 (Wood et al. 2019). For functional profiling, high-quality (filtered) reads were aligned against the SEED database (Overbeek et al. 2014) via translated homology search and annotated to Subsystems, or functional levels, 1-3 using Super-Focus (Silva et al. 2016). Functional gene abundance was expressed using pseudocounts (wherein counts of an individual functional gene are divided by the number of hits in the complete database) following the method described in Silva et al. (2016). In order to account for differences in sampling depth and composition between different samples, DESeq2 normalization (geometric means – Love et al. 2014) was used. To quantify differences in the distribution of taxa and functional genes between localities, pairwise dissimilarities were calculated using the Bray-Curtis index (Bray and Curtis 1957), and samples were compared using permutational multivariate analysis of variance (PERMANOVA – Anderson 2001), using functions from the vegan R package (Oksanen et al. 2024). Comparison of functional genes were made using the Level 3 SEED Subsystem (Overbeek et al. 2014). Alpha diversity of samples was quantified using the Shannon index (Shannon 1948) and compared between locations using the Kruskal-Wallis chi-squared test (Kruskal and Wallis 1952). Finally, to better understand the differences between these two localities, Wald tests (Wald 1943) were used to test for significant differences in the abundance of specific taxa and the expression of specific functional genes (using the DESeq2 R package – Love et al. 2014).
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,007 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,004 | 0,007 |
| Études des sciences et des technologies | 0,002 | 0,000 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,349 | 0,091 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».