Additional file 2 of Shared environments complicate the use of strain-resolved metagenomics to infer microbiome transmission
Notice bibliographique
Résumé
Supplementary Material 2. Supplementary figures: Figure S1. Representation of individuals in multiple dyads does not bias results. We iteratively downsampled the full dataset to include only a single randomly selected sample from each individual. The adjusted p-values as reported by a Tukey’s HSD Test were generally consistent with the results of the full dataset. Comparisons that were not significant in the full dataset (e.g., Not alive at same time – Different social groups) never rose to significance in our permutations, while comparisons that were significant in the full dataset (e.g., Not alive at same time – Longitudinal) typically remained significant despite the reduced sample size. Dashed line indicates padj=0.05. Figure S2. Mapping rates in the baboon gut metagenome dataset. (a) Database customization with metagenome-assembled genomes. Percentage of reads after quality filtering that aligned to microbial genomes in either the Unified Human Gastrointestinal Genome (UHGG) database or the UHGG database plus metagenome-assembled genomes from non-human primate fecal samples. (b) Mapping rates by month of sample collection. Mapping rates varied slightly by sampling month (ANOVA, p=0.053), with a tendency for samples from the dry season to map better than samples from the wet season. (c) Mapping rates by year of sample collection. Mapping rates did not vary by year. Figure S3. Bacterial taxa shared between donors and recipients of fecal microbiota transplants. (a) Family-level structure of bacterial taxa that were shared at the species level between subjects. Bar chart of species (aggregated into families) that exhibit species sharing between matched (left) or mismatched (right) pairs. Colors reflect different bacterial families and are scaled to represent the relative proportions of strain-sharing events in each category. (b) Species-level sharing of families among matched pairs relative to mismatched pairs. The log2 odds ratio represents the relative likelihood that bacterial families were shared at the species level between matched donor-recipient pairs, compared to their species sharing rates between mismatched pairs. (c) Strain-level sharing of families among matched pairs relative to mismatched pairs. The log2 odds ratio represents the relative likelihood that bacterial families were shared at the strain level between matched donor-recipient pairs, compared to their strain sharing rates between mismatched pairs. Error bars represent 95% confidence intervals. Error bars overlapping 0 (dashed line) are not significantly enriched or depleted in matched pairs. Figure S4. Taxonomic and functional characteristics of bacterial taxa in which strain sharing was detected between baboons in a wild population. (a) Family-level structure of bacterial taxa that were shared at the strain level between individuals. Bar chart of species (aggregated into families) that exhibit strain-sharing between pairs of baboons, based on a 99.999% ANI threshold. Colors reflect different bacterial families and are scaled to represent the relative proportions of strain-sharing events in each category. (b) Anaerobic metabolism across varying definitions of transmission. Strain sharing events were annotated based on phenotype information available in the Genomes OnLine database. (c) Population prevalence of bacterial species. Each strain sharing event was annotated according to the prevalence of the species to which it belonged, excluding later longitudinal samples to avoid double-counting individuals. The x-axis represents the number of baboons containing the species (of a maximum possible of 93). Figure S5. Diet and seasonality in Amboseli. (a) Dietary similarity between social groups. Dietary similarity was calculated using the Jaccard similarity index based on the dietary compositions of each pair of baboons that lived in different social groups at similar times, aggregated by month. (b) Variation in diets throughout the year. Colors reflect different foods consumed by the baboons between 2007-2017 and are scaled to represent the relative proportions consumed in each month. Figure S6. A consensus-based approach does not detect an effect of rainfall on strain sharing. Rainfall was measured daily using a rain gauge and summed by month for samples taken from baboons that lived at different times. The y-axis represents strain sharing based on the consensus average nucleotide identity (conANI) value calculated by inStrain. This metric considers two genomes to differ at a given site if their consensus alleles are different (i.e., it ignores minor allele sharing). (b) Microdiversity-aware approach (reproduced from Figure 4b). This metric considers two genomes to differ at a given site if they do not share any alleles at that site (major or minor). Figure S7. Host characteristics that did not predict strain sharing rates. (a) Number of years between samples. Plot includes only baboon pairs that lived at different times, as the remaining dyad types were intentionally sampled within short time periods. (b) Difference in ages at times of sampling. Chronological ages were known with high confidence because all subjects were born in regularly censused study groups. Figure S8. Strain sharing analysis of FMT dataset with StrainPhlAn pipeline. (a) Correlation between inStrain and StrainPhlAn estimates of strain sharing. Each point represents a donor-recipient pair (either matched or mismatched) in the FMT dataset. Samples with fewer than three shared species (i.e., the denominator in the calculation of strain sharing rates) are excluded from this visualization. (b) Strain sharing across varying definitions of transmission. The percentage of strain sharing events among matched donor-recipient dyads (blue) and all other comparisons (red) is shown following serially more stringent filtering criteria. Asterisks represent significant differences between matched and mismatched cohorts based on t-test and Benjamini-Hochberg correction: (***) p<0.001; (**) 0.001≤p<0.01; (*) 0.01≤p<0.05. Figure S9. Strain sharing analysis of baboon dataset with StrainPhlAn pipeline. (a) Correlation between inStrain and StrainPhlAn estimates of strain sharing. Each point represents a donor-recipient pair in the FMT dataset. Samples with fewer than three shared species (i.e., the denominator in the calculation of strain sharing rates) are excluded from this visualization. (b) Strain sharing rates across dyad types. The percentage of shared strains between each dyad (≤0.1 normalized phylogenetic distance).
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,028 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,003 | 0,005 |
| Études des sciences et des technologies | 0,002 | 0,001 |
| Communication savante | 0,004 | 0,004 |
| Science ouverte | 0,003 | 0,003 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,775 | 0,170 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».