Evolutionary criteria outperform operational approaches in producing ecologically relevant fungal species inventories
Bibliographic record
Abstract
Analyses of the structure and function of microbial communities are highly constrained by the diversity of organisms present within most environmental samples. A common approach is to rely almost entirely on DNA sequence data for estimates of microbial diversity, but to date there is no objective method of clustering sequences into groups that is grounded in evolutionary theory of what constitutes a biological lineage. The general mixed Yule-coalescent (GMYC) model uses a likelihood-based approach to distinguish population-level processes within lineages from processes associated with speciation and extinction, thus identifying a distinct point where extant lineages became independent. Using two independent surveys of DNA sequences associated with a group of ubiquitous plant-symbiotic fungi, we compared estimates of species richness derived using the GMYC model to those based on operational taxonomic units (OTUs) defined by fixed levels of sequence similarity. The model predicted lower species richness in these surveys than did traditional methods of sequence similarity. Here, we show for the first time that groups delineated by the GMYC model better explained variation in the distribution of fungi in relation to putative niche-based variables associated with host species identity, edaphic factors, and aspects of how the sampled ecosystems were managed. Our results suggest the coalescent-based GMYC model successfully groups environmental sequences of fungi into clusters that are ecologically more meaningful than more arbitrary approaches for estimating species richness.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.017 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".