Measuring Branch Support in Species Trees Obtained by Gene Tree Parsimony
Bibliographic record
Abstract
Several methods have recently been developed that allow the reconstruction of species trees from gene trees, an important achievement in our ongoing quest to obtain reliable species phylogenies. However, considerably less attention has been given to evaluating the accuracy of species trees' estimates. Four methods for measuring branch support of species trees are tested in this study in a gene tree parsimony framework: 1) bootstrap lineages (BL) (sequences) within species, 2) bootstrap characters (BC) within genes (i.e., the standard nonparametric bootstrap), 3) bootstrap lineages and characters (BLC), and 4) posterior probability gene tree sampling (PPGTS) (where, for each resampled data set, gene trees are sampled according to their posterior probability). For each method, n species trees are reconstructed from n resampled data sets and the branch support consists in the percentage of the n species trees in which a branch is recovered. The 4 methods were tested for several species trees and for different sampling efforts (i.e., number of genes and individuals sampled) using coalescent simulations. PPGTS performed best overall with lowest Type I and II error rates, followed by BLC. The BL and BC methods had higher error rates. This suggests that in order to properly measure branch support in a species tree context, it is important to account for the uncertainty involved in reconstructing gene trees from DNA sequences as well as that involved in reconstructing the species tree from individual gene trees. With the parameters used in the simulations, sampling more individuals per species resulted in similar improvements in support values as when sampling more genes. Moreover, sampling more individuals per species appeared to be important for escaping the anomaly zone present when only 1 sequence was sampled. We also apply the 4 methods to obtain branch supports for the species phylogeny of diploid wild roses (Rosa) in North America.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.064 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".