MétaCan
Menu
← Back to cohort

Additional file 2 of Shared environments complicate the use of strain-resolved metagenomics to infer microbiome transmission

2025· article· en· W6921029190 on OpenAlexaff

Bibliographic record

VenueFigshare · 2025
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGut microbiota and health
Canadian institutionsCanadian Institute for Advanced Research
Fundersnot available
KeywordsMetagenomicsMicrobiomeGenomeSample (material)TaxonPedigree chart

Abstract

fetched live from OpenAlex

Supplementary Material 2. Supplementary figures: Figure S1. Representation of individuals in multiple dyads does not bias results. We iteratively downsampled the full dataset to include only a single randomly selected sample from each individual. The adjusted p-values as reported by a Tukey’s HSD Test were generally consistent with the results of the full dataset. Comparisons that were not significant in the full dataset (e.g., Not alive at same time – Different social groups) never rose to significance in our permutations, while comparisons that were significant in the full dataset (e.g., Not alive at same time – Longitudinal) typically remained significant despite the reduced sample size. Dashed line indicates padj=0.05. Figure S2. Mapping rates in the baboon gut metagenome dataset. (a) Database customization with metagenome-assembled genomes. Percentage of reads after quality filtering that aligned to microbial genomes in either the Unified Human Gastrointestinal Genome (UHGG) database or the UHGG database plus metagenome-assembled genomes from non-human primate fecal samples. (b) Mapping rates by month of sample collection. Mapping rates varied slightly by sampling month (ANOVA, p=0.053), with a tendency for samples from the dry season to map better than samples from the wet season. (c) Mapping rates by year of sample collection. Mapping rates did not vary by year. Figure S3. Bacterial taxa shared between donors and recipients of fecal microbiota transplants. (a) Family-level structure of bacterial taxa that were shared at the species level between subjects. Bar chart of species (aggregated into families) that exhibit species sharing between matched (left) or mismatched (right) pairs. Colors reflect different bacterial families and are scaled to represent the relative proportions of strain-sharing events in each category. (b) Species-level sharing of families among matched pairs relative to mismatched pairs. The log2 odds ratio represents the relative likelihood that bacterial families were shared at the species level between matched donor-recipient pairs, compared to their species sharing rates between mismatched pairs. (c) Strain-level sharing of families among matched pairs relative to mismatched pairs. The log2 odds ratio represents the relative likelihood that bacterial families were shared at the strain level between matched donor-recipient pairs, compared to their strain sharing rates between mismatched pairs. Error bars represent 95% confidence intervals. Error bars overlapping 0 (dashed line) are not significantly enriched or depleted in matched pairs. Figure S4. Taxonomic and functional characteristics of bacterial taxa in which strain sharing was detected between baboons in a wild population. (a) Family-level structure of bacterial taxa that were shared at the strain level between individuals. Bar chart of species (aggregated into families) that exhibit strain-sharing between pairs of baboons, based on a 99.999% ANI threshold. Colors reflect different bacterial families and are scaled to represent the relative proportions of strain-sharing events in each category. (b) Anaerobic metabolism across varying definitions of transmission. Strain sharing events were annotated based on phenotype information available in the Genomes OnLine database. (c) Population prevalence of bacterial species. Each strain sharing event was annotated according to the prevalence of the species to which it belonged, excluding later longitudinal samples to avoid double-counting individuals. The x-axis represents the number of baboons containing the species (of a maximum possible of 93). Figure S5. Diet and seasonality in Amboseli. (a) Dietary similarity between social groups. Dietary similarity was calculated using the Jaccard similarity index based on the dietary compositions of each pair of baboons that lived in different social groups at similar times, aggregated by month. (b) Variation in diets throughout the year. Colors reflect different foods consumed by the baboons between 2007-2017 and are scaled to represent the relative proportions consumed in each month. Figure S6. A consensus-based approach does not detect an effect of rainfall on strain sharing. Rainfall was measured daily using a rain gauge and summed by month for samples taken from baboons that lived at different times. The y-axis represents strain sharing based on the consensus average nucleotide identity (conANI) value calculated by inStrain. This metric considers two genomes to differ at a given site if their consensus alleles are different (i.e., it ignores minor allele sharing). (b) Microdiversity-aware approach (reproduced from Figure 4b). This metric considers two genomes to differ at a given site if they do not share any alleles at that site (major or minor). Figure S7. Host characteristics that did not predict strain sharing rates. (a) Number of years between samples. Plot includes only baboon pairs that lived at different times, as the remaining dyad types were intentionally sampled within short time periods. (b) Difference in ages at times of sampling. Chronological ages were known with high confidence because all subjects were born in regularly censused study groups. Figure S8. Strain sharing analysis of FMT dataset with StrainPhlAn pipeline. (a) Correlation between inStrain and StrainPhlAn estimates of strain sharing. Each point represents a donor-recipient pair (either matched or mismatched) in the FMT dataset. Samples with fewer than three shared species (i.e., the denominator in the calculation of strain sharing rates) are excluded from this visualization. (b) Strain sharing across varying definitions of transmission. The percentage of strain sharing events among matched donor-recipient dyads (blue) and all other comparisons (red) is shown following serially more stringent filtering criteria. Asterisks represent significant differences between matched and mismatched cohorts based on t-test and Benjamini-Hochberg correction: (***) p<0.001; (**) 0.001≤p<0.01; (*) 0.01≤p<0.05. Figure S9. Strain sharing analysis of baboon dataset with StrainPhlAn pipeline. (a) Correlation between inStrain and StrainPhlAn estimates of strain sharing. Each point represents a donor-recipient pair in the FMT dataset. Samples with fewer than three shared species (i.e., the denominator in the calculation of strain sharing rates) are excluded from this visualization. (b) Strain sharing rates across dyad types. The percentage of shared strains between each dyad (≤0.1 normalized phylogenetic distance).

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.028
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.775
Threshold uncertainty score0.320

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.028
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0030.005
Science and technology studies0.0020.001
Scholarly communication0.0040.004
Open science0.0030.003
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.7750.170

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.045
GPT teacher head0.259
Teacher spread0.214 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueFigshare→Same topicGut microbiota and health→French-language works237,207→