MétaCan
Menu
← Retour à la cohorte
Enregistrement W4394560351 · doi:10.6084/m9.figshare.22615279

Additional file 1 of Degradation pathways for organic matter of terrestrial origin are widespread and expressed in Arctic Ocean microbiomes

2023· dataset· en· W4394560351 sur OpenAlexaff
Thomas Grevesse, Céline Guéguen, Vera E. Onana, David A. Walsh

Notice bibliographique

RevueFigshare · 2023
Typedataset
Langueen
DomaineEnvironmental Science
ThématiqueMicrobial Community Ecology and Physiology
Établissements canadiensUniversité de SherbrookeConcordia University
Organismes subventionnairesnon disponible
Mots-clésArcticMicrobiomeDegradation (telecommunications)Organic matterThe arcticOceanographyEarth scienceEnvironmental scienceGeographyBiologyEcologyGeologyComputer scienceBioinformatics

Résumé

récupéré en direct d'OpenAlex

Additional file 1: Figure S1. Estimation of the rank for the NMF analysis of the EC abundance matrices annotated from metagenomes (top panels) and metatranscriptomes (bottom panels). Left panels: evolution of various parameters as a function of the rank used in the NMF analysis. Cophenetic correlation represents the correlation between the sample distances from the consensus matrix and the cophenetic distance between these samples when they are clustered. The rss is the residual sum of squares between the original EC abundance matrix and its estimate using the NMF algorithm. The dispersion is defined as 1-rss/Σi,j (Vi,j)2 (Vi,j are the entries of the EC abundance matrix) and estimates the fraction of variance of the EC abundance matrix explained by the NMF results. Residuals is the sum of residuals between the original EC abundance matrix and the matrix estimated using the NMF. Right panels: consensus matrices based on clustering the coefficient matrices at each of the 100 runs of the NMF analysis. The heatmap represents the fraction of times 2 samples fall in the same clusters out of 100 runs. Figure S2. Heatmaps of the basis matrix (left) and coefficient matrix (right) obtained after running an NMF analysis on the EC abundance matrix annotated from metagenomes and using a rank value of 4. Figure S3. Heatmaps of the basis matrix (left) and coefficient matrix (right) obtained after running an NMF analysis on the EC abundance matrix annotated from metatranscriptomes and using a rank value of 4. Figure S4. Lignin-derived aromatic compound degradation pathways completeness in the 4 water column features (surface, SCM, FDOMmax, and deep water) for metagenomes (left) and metatranscriptomes (right). Figure S5. Normalized abundance per water column feature of KO number markers of aromatic compounds degradation pathways annotated from metagenomes. Figure S6. Normalized abundance per water column feature of KO numbers markers of aromatic compounds degradation pathways annotated from metatranscriptomes. Figure S7. Estimated fraction of the microbiome harboring genes annotated with KO marker of aromatic compounds. Figure S8. Completeness and contamination of the set of 1772 MAGs reconstructed from 22 individual metagenomes. The vertical and horizontal dotted lines represent 10% contamination and 30% completeness respectively. Selected MAGs are the 46 MAGs that have been selected as most implicated in lignin-derived aromatic compound degradation based on the number and completeness of their aromatic compound degradation pathways. Figure S9. Taxonomic identity of the MAGs harboring genes annotated with KO marker of aromatic compounds degradation pathways. The taxonomy is displayed at the class level. Figure S10. Phylogenetic tree reconstructed from the concatenation of 120 conserved genes for the bacterial genomes of the MAGs dataset. The heatmap represents the completeness of the lignin-derived aromatic compounds degradation funneling pathways. Red stars correspond to the MAGs selected based on the amount and completeness of pathways they harbor. Figure S11. Taxonomic identity of the MAGs most implicated in the degradation of lignin-derived aromatic compounds. Taxonomy is displayed at the order level and based on the tree placement of the MAGs, using the GTDB. Figure S12. Heatmap of the average nucleotide identity for the 46 MAGs most implicated in the degradation of lignin-derived aromatic compounds.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,019
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesCharge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Jeu de données · Signal consensuel: Jeu de données
Score de désaccord entre enseignants0,866
Score d'incertitude au seuil0,191

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0020,019
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0020,002
Bibliométrie0,0030,006
Études des sciences et des technologies0,0020,001
Communication savante0,0020,003
Science ouverte0,0030,002
Intégrité de la recherche0,0020,002
Charge utile insuffisante (le modèle a refusé de juger)0,8660,154

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,034
Tête enseignante GPT0,235
Écart entre enseignants0,201 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeSans objet
Domainenon disponible
GenreJeu de données

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2023
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueFigshare→Même sujetMicrobial Community Ecology and Physiology→Travaux en français237 207→